← DatVietVAC
Senior Site Reliability Engineer SRE, GCP, Kubernetes
DatVietVAC · Hồ Chí Minh
Ứng tuyển tại trang chính thức ↗
Loại hình
Toàn thời gian
Hình thức
Tại văn phòng
Cấp bậc
Nhân viên
Ngành nghề
Kỹ thuật / Cơ khí
Mức lương
Thương lượng
Địa điểm
Quận 3, Hồ Chí Minh, Hồ Chí Minh
Tổng quan
- Cloud Infrastructure Architecture
- Design and implement multi-environment infrastructure, including Production, Staging, and Development, using Google Cloud services such as GKE, Cloud SQL, Memorystore, and Pub/Sub.
- Design the hybrid-cloud architecture between Google Cloud and CMC Cloud, including private connectivity, VPN configuration, network segmentation, and data-flow separation.
- Build and manage infrastructure as code using Terraform.
- Develop and maintain CI/CD pipelines and automation tools that support engineering teams throughout the software development and release lifecycle.
- Establish a comprehensive observability platform covering metrics, logs, traces, dashboards, and alerting.
- Ensure that the infrastructure architecture supports scalability, maintainability, security, and local data-residency requirements.
- Site Reliability and Incident Response
- Define and manage Service Level Indicators, Service Level Objectives, and error budgets for core platform services.
- Take ownership of platform availability, reliability, scalability, and operational readiness.
- Lead capacity planning and load testing for high-traffic product launches, ticket or merchandise on-sales, and live events.
- Design autoscaling strategies, traffic-management mechanisms, and overload-protection measures.
- Participate in the 24/7 production on-call rotation and lead the response to high-severity incidents.
- Lead blameless postmortems, identify root causes, and ensure that corrective and preventive actions are completed.
- Develop incident-response procedures, operational runbooks, and disaster-recovery plans.
- Coach and mentor engineers in production operations, incident management, and reliability practices.
- Security and Compliance
- Implement infrastructure security controls, including WAF, DDoS protection, and bot-management solutions using tools such as Cloudflare and Google Cloud Armor.
- Manage secrets, encryption, identity, and access controls across cloud environments.
- Establish backup, recovery, and disaster-recovery processes and conduct periodic recovery drills.
- Collaborate with relevant teams on penetration testing, vulnerability remediation, security audits, and compliance requirements.
- Ensure appropriate monitoring and audit logging for sensitive infrastructure and operational activities.
- Performance and Cost Optimization
- Monitor, analyze, and optimize cloud infrastructure costs through rightsizing, committed-use discounts, storage lifecycle management, and egress optimization.
- Investigate performance bottlenecks and implement solutions to improve platform efficiency, scalability, and resilience.
- Provide infrastructure and reliability recommendations to the Project Lead and engineering teams.
- Evaluate and propose technologies that support the platform’s long-term architecture and business requirements.
- Bachelor’s degree or higher in Information Technology, Software Engineering, Computer Science, or a related field.
- At least five years of experience as a DevOps Engineer, Site Reliability Engineer, System Engineer, Cloud Engineer, or in a similar infrastructure role.
- At least one year of experience at an equivalent senior level or in a technical leadership role.
- Proven experience designing or operating production systems with high traffic or significant peak-load events, such as flash sales, on-sales, live events, entertainment platforms, or e-commerce platforms.
- Hands-on experience operating Kubernetes and cloud infrastructure in a production environment.
- Production experience with Google Cloud Platform is mandatory.
- Advanced knowledge of networking and security concepts, including TCP/IP, HTTP/1.1, HTTP/2, HTTP/3, DNS, gRPC, VPC peering, Cloud VPN, and on-premises-to-cloud connectivity.
- Strong understanding of distributed systems, microservices, clustering, replication, failover, load balancing, and autoscaling.
- Strong hands-on experience with Kubernetes, particularly Google Kubernetes Engine, and container technologies in production environments.
- Proficiency in infrastructure as code and automation tools, particularly Terraform and Ansible.
- Strong experience building and maintaining CI/CD pipelines using GitHub Actions, Jenkins, or equivalent tools.
- Strong Linux administration skills, including experience with Ubuntu or CentOS.
- Hands-on experience with Google Cloud services; additional AWS or Azure experience is an advantage.
Yêu cầu
- Willingness to participate in the production on-call rotation and support major on-sale periods or live events when required.
- Ability to work under pressure during critical releases, peak-traffic events, and production incidents.
- Why You'll Love Working Here
- Full statutory insurance, including Social Insurance, Health Insurance and Unemployment Insurance, based on 100% of the official salary and in compliance with Vietnamese labor regulations.
- Working hours: Monday to Friday, from 8:30 AM to 5:30 PM, with a one-hour lunch break.
- 14 days of annual leave.
- Company-provided working equipment, including a laptop or desktop computer.
- Employee parking area.
- Annual PMP performance bonus, subject to individual KPI achievement and the Company’s business performance.
- Periodic health check-ups.
- Employee engagement programs and internal activities throughout the year.
Quyền lợi
Chế độ thưởngLaptop / Thiết bịGửi xe
Tóm tắt thông tin từ tin tuyển dụng chính thức. Xem bản gốc ↗
Quan tâm đến vị trí này?
Bạn sẽ được chuyển đến trang ứng tuyển chính thức của nhà tuyển dụng.
Ứng tuyển tại trang chính thức ↗
Đây là doanh nghiệp của bạn?
Nhận quản lý trang, yêu cầu chỉnh sửa hoặc gỡ bỏ
→