← DatVietVAC
SRE DevOps, GCP, Kubernetes, Terraform
DatVietVAC · Hồ Chí Minh
Apply on official site ↗
Type
Full-time
Work mode
On-site
Level
Staff
Industry
Other
Salary
Thương lượng
Location
Quận 3, Hồ Chí Minh, Hồ Chí Minh
Overview
- Cloud Infrastructure and Platform Operations
- Operate and maintain application infrastructure in cloud environments, including compute, databases, caching, message queues, and object storage.
- Manage and continuously improve Development, Staging, and Production environments.
- Operate and improve CI/CD pipelines that support application development, testing, and deployment.
- Provision and manage infrastructure as code using Terraform and Ansible.
- Manage network configurations, firewalls, routing, VPC connectivity, and cloud VPNs.
- Operate containerized workloads and Kubernetes clusters, particularly Google Kubernetes Engine.
- Perform routine system upgrades, security patching, configuration changes, and infrastructure maintenance.
- Maintain infrastructure documentation and configuration standards.
- Monitoring, Reliability, and Incident Response
- Monitor platform availability, performance, capacity, and reliability against agreed service-level objectives.
- Build and maintain monitoring, logging, dashboard, and alerting systems using Prometheus, Grafana, Google Cloud Operations, Datadog, or equivalent tools.
- Participate in the production on-call rotation and respond to incidents in a timely manner.
- Troubleshoot application, infrastructure, network, database, cache, and message-queue issues.
- Support root-cause analysis and contribute to post-incident reports and corrective actions.
- Develop and maintain operational runbooks for platform services and recurring incident scenarios.
- Analyze operational data and system metrics to identify reliability and performance improvements.
- Automation and Continuous Improvement
- Develop automation tools and scripts to reduce manual operational work and repetitive engineering tasks.
- Improve deployment automation, environment consistency, and release reliability.
- Support load and performance testing before high-traffic on-sale periods and major events.
- Assist with capacity preparation, scaling configuration, and infrastructure readiness.
- Identify technical debt, operational risks, and opportunities to improve platform resilience.
- Support backup, recovery, and system-security activities.
- Contribute to continuous improvements in infrastructure management, monitoring, and incident response.
- Cross-functional Collaboration
- Collaborate closely with Backend Engineers, AI-Native Engineers, QC Engineers, and Product Owners to improve development and release processes.
- Support engineering teams in resolving environment, deployment, configuration, and production issues.
- Participate in release planning and production-readiness reviews.
- Communicate technical issues, operational risks, and dependencies clearly to relevant stakeholders.
- Work independently when required while contributing to the overall effectiveness of the engineering team.
- Maintain technical documentation, operational procedures, and service runbooks.
- Bachelor’s degree in information technology, Software Engineering, Computer Science, or a related field.
- At least three years of experience as a DevOps Engineer, Site Reliability Engineer, System Engineer, Cloud Engineer, or in a related role.
- Hands-on experience operating cloud infrastructure and production systems.
- Experience with containerized applications and Kubernetes in a production environment.
- Experience provisioning and managing infrastructure through Infrastructure as Code.
- Experience supporting production deployments and responding to system incidents.
- Experience with transactional, e-commerce, OTT, entertainment, or other high-traffic platforms is preferred.
- Solid knowledge of TCP/IP, HTTP/1.1, HTTP/2, DNS, gRPC, VPC peering, Cloud VPN, and connectivity between cloud environments.
Requirements
- Willingness to participate in the production on-call rotation.
- Willingness to provide operational support during major releases, on-sale periods, or live events when required.
- Why You'll Love Working Here
- Full statutory insurance, including Social Insurance, Health Insurance and Unemployment Insurance, based on 100% of the official salary and in compliance with Vietnamese labor regulations.
- Working hours: Monday to Friday, from 8:30 AM to 5:30 PM, with a one-hour lunch break.
- 14 days of annual leave.
- Company-provided working equipment, including a laptop or desktop computer.
- Employee parking area.
- Annual PMP performance bonus, subject to individual KPI achievement and the Company’s business performance.
- Periodic health check-ups.
- Employee engagement programs and internal activities throughout the year.
Benefits
BonusLaptop / equipmentParking
Summary of facts from the official posting. View original ↗
Interested in this role?
You'll be taken to the employer's official application page.
Apply on official site ↗
Is this your business?
Claim this page, request edits or removal
→