Đến nội dung chính
Tuyển Dụng
← DatVietVAC

Senior Site Reliability Engineer SRE, GCP, Kubernetes

DatVietVAC · Hồ Chí Minh
Ứng tuyển tại trang chính thức ↗
Loại hình
Toàn thời gian
Hình thức
Tại văn phòng
Cấp bậc
Nhân viên
Ngành nghề
Kỹ thuật / Cơ khí
Mức lương
Thương lượng
Địa điểm
Quận 3, Hồ Chí Minh, Hồ Chí Minh

Tổng quan

  • Cloud Infrastructure Architecture
  • Design and implement multi-environment infrastructure, including Production, Staging, and Development, using Google Cloud services such as GKE, Cloud SQL, Memorystore, and Pub/Sub.
  • Design the hybrid-cloud architecture between Google Cloud and CMC Cloud, including private connectivity, VPN configuration, network segmentation, and data-flow separation.
  • Build and manage infrastructure as code using Terraform.
  • Develop and maintain CI/CD pipelines and automation tools that support engineering teams throughout the software development and release lifecycle.
  • Establish a comprehensive observability platform covering metrics, logs, traces, dashboards, and alerting.
  • Ensure that the infrastructure architecture supports scalability, maintainability, security, and local data-residency requirements.
  • Site Reliability and Incident Response
  • Define and manage Service Level Indicators, Service Level Objectives, and error budgets for core platform services.
  • Take ownership of platform availability, reliability, scalability, and operational readiness.
  • Lead capacity planning and load testing for high-traffic product launches, ticket or merchandise on-sales, and live events.
  • Design autoscaling strategies, traffic-management mechanisms, and overload-protection measures.
  • Participate in the 24/7 production on-call rotation and lead the response to high-severity incidents.
  • Lead blameless postmortems, identify root causes, and ensure that corrective and preventive actions are completed.
  • Develop incident-response procedures, operational runbooks, and disaster-recovery plans.
  • Coach and mentor engineers in production operations, incident management, and reliability practices.
  • Security and Compliance
  • Implement infrastructure security controls, including WAF, DDoS protection, and bot-management solutions using tools such as Cloudflare and Google Cloud Armor.
  • Manage secrets, encryption, identity, and access controls across cloud environments.
  • Establish backup, recovery, and disaster-recovery processes and conduct periodic recovery drills.
  • Collaborate with relevant teams on penetration testing, vulnerability remediation, security audits, and compliance requirements.
  • Ensure appropriate monitoring and audit logging for sensitive infrastructure and operational activities.
  • Performance and Cost Optimization
  • Monitor, analyze, and optimize cloud infrastructure costs through rightsizing, committed-use discounts, storage lifecycle management, and egress optimization.
  • Investigate performance bottlenecks and implement solutions to improve platform efficiency, scalability, and resilience.
  • Provide infrastructure and reliability recommendations to the Project Lead and engineering teams.
  • Evaluate and propose technologies that support the platform’s long-term architecture and business requirements.
  • Bachelor’s degree or higher in Information Technology, Software Engineering, Computer Science, or a related field.
  • At least five years of experience as a DevOps Engineer, Site Reliability Engineer, System Engineer, Cloud Engineer, or in a similar infrastructure role.
  • At least one year of experience at an equivalent senior level or in a technical leadership role.
  • Proven experience designing or operating production systems with high traffic or significant peak-load events, such as flash sales, on-sales, live events, entertainment platforms, or e-commerce platforms.
  • Hands-on experience operating Kubernetes and cloud infrastructure in a production environment.
  • Production experience with Google Cloud Platform is mandatory.
  • Advanced knowledge of networking and security concepts, including TCP/IP, HTTP/1.1, HTTP/2, HTTP/3, DNS, gRPC, VPC peering, Cloud VPN, and on-premises-to-cloud connectivity.
  • Strong understanding of distributed systems, microservices, clustering, replication, failover, load balancing, and autoscaling.
  • Strong hands-on experience with Kubernetes, particularly Google Kubernetes Engine, and container technologies in production environments.
  • Proficiency in infrastructure as code and automation tools, particularly Terraform and Ansible.
  • Strong experience building and maintaining CI/CD pipelines using GitHub Actions, Jenkins, or equivalent tools.
  • Strong Linux administration skills, including experience with Ubuntu or CentOS.
  • Hands-on experience with Google Cloud services; additional AWS or Azure experience is an advantage.

Yêu cầu

  • Willingness to participate in the production on-call rotation and support major on-sale periods or live events when required.
  • Ability to work under pressure during critical releases, peak-traffic events, and production incidents.
  • Why You'll Love Working Here
  • Full statutory insurance, including Social Insurance, Health Insurance and Unemployment Insurance, based on 100% of the official salary and in compliance with Vietnamese labor regulations.
  • Working hours: Monday to Friday, from 8:30 AM to 5:30 PM, with a one-hour lunch break.
  • 14 days of annual leave.
  • Company-provided working equipment, including a laptop or desktop computer.
  • Employee parking area.
  • Annual PMP performance bonus, subject to individual KPI achievement and the Company’s business performance.
  • Periodic health check-ups.
  • Employee engagement programs and internal activities throughout the year.

Quyền lợi

Chế độ thưởngLaptop / Thiết bịGửi xe

Tóm tắt thông tin từ tin tuyển dụng chính thức. Xem bản gốc ↗

Quan tâm đến vị trí này?

Bạn sẽ được chuyển đến trang ứng tuyển chính thức của nhà tuyển dụng.

Ứng tuyển tại trang chính thức ↗
Đây là doanh nghiệp của bạn? Nhận quản lý trang, yêu cầu chỉnh sửa hoặc gỡ bỏ