Skip to content
Tuyển Dụng
← CarDoctor

Infrastructure Lead Cloud

CarDoctor · Hà Nội
Apply on official site ↗
Type
Full-time
Work mode
On-site
Level
Staff
Industry
Other
Salary
Thương lượng
Location
Quận Nam Từ Liêm, Hà Nội, Hà Nội

Overview

  • Working Location: 5BT2, Me Tri Ha Urban Area, Tu Liem Ward, Hanoi
  • Working Hours: 08:00 AM – 05:30 PM | Monday – Friday
  • Lead the design, operation, and continuous improvement of the company's hybrid infrastructure, spanning both cloud (AWS/GCP) and on-premise (data center/colocation) environments.
  • Ensure high availability, security, performance, and cost-efficiency across both cloud and on-prem systems, with a consistent architecture and operating model between the two.
  • Lead the Infra team (~4-6 engineers) and coordinate with internal teams and vendors to ensure reliable infrastructure operations end-to-end.
  • A. Infrastructure Architecture & Capacity Planning
  • Design and own the hybrid infrastructure architecture: cloud-native workloads on AWS/GCP (EKS, EC2, S3, RDS, VPC, ALB/NLB...) and on-premise infrastructure (data center/colocation: compute, storage, network, security); standardize dev, staging, and production environments across both.
  • Lead capacity planning for both cloud and on-prem, based on user growth, tenant/station count, and workload trends; build a 12–24 month hardware/cloud investment roadmap.
  • Design High Availability and Disaster Recovery across environments (multi-AZ/multi-site, active-active or active-standby); define RPO/RTO per service tier and run periodic DR drills.
  • Decide, workload by workload, what runs on cloud vs. on-prem, based on cost, performance, compliance, and Vietnam data-residency requirements.
  • B. Platform Operations, CI/CD & Container Orchestration
  • Lead Kubernetes operations at scale across both environments (EKS and on-prem K8s) using Helm, with autoscaling/self-healing; support multi-tenant architecture for platform products (e.g. Service Station Portal).
  • Build and standardize CI/CD (Jenkins/GitLab CI/ArgoCD) and IaC (Terraform, Ansible, Helm) so releases to either environment are safe, fast, and consistent.
  • Operate the data and middleware layer (PostgreSQL/MySQL clusters, Redis, Kafka/RabbitMQ, Elasticsearch/OpenSearch, object storage) and real-time infrastructure (VoIP/SIP, WebRTC, WebSocket) where applicable; ensure performance during peak periods.
  • C. Observability, Security & Cost Efficiency
  • Build and maintain observability (Prometheus, Grafana, ELK/Loki, CloudWatch) across cloud and on-prem; define and enforce SLIs/SLOs/SLAs, and lead incident response and blameless RCA.
  • Enforce least-privilege IAM, network/firewall security, and secret management across both environments; integrate security monitoring into reliability workflows.
  • Apply FinOps/cost-efficiency practices on cloud (rightsizing, autoscaling, Savings Plans) and cost-aware capacity planning on-prem; balance reliability vs. cost trade-offs across the hybrid estate.
  • D. Team Leadership & Collaboration
  • Manage, mentor, and develop the Infra team (~4-6 engineers); own the infrastructure technical roadmap across cloud and on-prem.
  • Collaborate with Backend, QA, and Security teams to standardize release, rollback, and versioning processes; report infrastructure status, incidents, and cost to the CTO/Management regularly.
  • Your Skills and Experience

Requirements

  • Minimum 7 years of experience in DevOps/SRE/Infrastructure, including experience leading a team; experience spanning both cloud and on-premise environments is strongly preferred.
  • Proficient with core AWS services (EKS, EC2, S3, RDS, IAM, VPC, ALB/NLB, CloudWatch) and Terraform (required).
  • Hands-on experience managing on-premise infrastructure: Servers, Network, Storage, Virtualization (VMware/Hyper-V), Backup/DR.
  • Strong understanding of Kubernetes operations at scale (EKS and/or on-prem), Helm, and multi-tenant architecture.
  • Experience with CI/CD (Jenkins/GitLab CI/ArgoCD) and implementing observability (Prometheus/Grafana/ELK); experience defining SLIs/SLOs/SLAs.
  • Solid understanding of system security (IAM/RBAC, secret management, least privilege, network/firewall) and FinOps/cost optimization.
  • Experience operating database clusters, caching, message queues, search, or object storage (PostgreSQL/MySQL, Redis, Kafka/RabbitMQ, Elasticsearch, Ceph/MinIO) is a plus.
  • Proven ability to lead/mentor a technical team, with strong communication skills across technical teams, senior stakeholders, and vendors.
  • Why You'll Love Working Here

Benefits

  • Competitive salary package based on qualifications and experience.
BonusCompany tripsTrainingLaptop / equipment

Summary of facts from the official posting. View original ↗

Interested in this role?

You'll be taken to the employer's official application page.

Apply on official site ↗
Is this your business? Claim this page, request edits or removal
→