Type
Full-time
Work mode
On-site
Level
Staff
Industry
Other
Salary
Thương lượng
Location
Quận Nam Từ Liêm, Hà Nội, Hà Nội
Overview
- Working Location: 5BT2, Me Tri Ha Urban Area, Tu Liem Ward, Hanoi
- Working Hours: 08:00 AM – 05:30 PM | Monday – Friday
- Lead the design, operation, and continuous improvement of the company's hybrid infrastructure, spanning both cloud (AWS/GCP) and on-premise (data center/colocation) environments.
- Ensure high availability, security, performance, and cost-efficiency across both cloud and on-prem systems, with a consistent architecture and operating model between the two.
- Lead the Infra team (~4-6 engineers) and coordinate with internal teams and vendors to ensure reliable infrastructure operations end-to-end.
- A. Infrastructure Architecture & Capacity Planning
- Design and own the hybrid infrastructure architecture: cloud-native workloads on AWS/GCP (EKS, EC2, S3, RDS, VPC, ALB/NLB...) and on-premise infrastructure (data center/colocation: compute, storage, network, security); standardize dev, staging, and production environments across both.
- Lead capacity planning for both cloud and on-prem, based on user growth, tenant/station count, and workload trends; build a 12–24 month hardware/cloud investment roadmap.
- Design High Availability and Disaster Recovery across environments (multi-AZ/multi-site, active-active or active-standby); define RPO/RTO per service tier and run periodic DR drills.
- Decide, workload by workload, what runs on cloud vs. on-prem, based on cost, performance, compliance, and Vietnam data-residency requirements.
- B. Platform Operations, CI/CD & Container Orchestration
- Lead Kubernetes operations at scale across both environments (EKS and on-prem K8s) using Helm, with autoscaling/self-healing; support multi-tenant architecture for platform products (e.g. Service Station Portal).
- Build and standardize CI/CD (Jenkins/GitLab CI/ArgoCD) and IaC (Terraform, Ansible, Helm) so releases to either environment are safe, fast, and consistent.
- Operate the data and middleware layer (PostgreSQL/MySQL clusters, Redis, Kafka/RabbitMQ, Elasticsearch/OpenSearch, object storage) and real-time infrastructure (VoIP/SIP, WebRTC, WebSocket) where applicable; ensure performance during peak periods.
- C. Observability, Security & Cost Efficiency
- Build and maintain observability (Prometheus, Grafana, ELK/Loki, CloudWatch) across cloud and on-prem; define and enforce SLIs/SLOs/SLAs, and lead incident response and blameless RCA.
- Enforce least-privilege IAM, network/firewall security, and secret management across both environments; integrate security monitoring into reliability workflows.
- Apply FinOps/cost-efficiency practices on cloud (rightsizing, autoscaling, Savings Plans) and cost-aware capacity planning on-prem; balance reliability vs. cost trade-offs across the hybrid estate.
- D. Team Leadership & Collaboration
- Manage, mentor, and develop the Infra team (~4-6 engineers); own the infrastructure technical roadmap across cloud and on-prem.
- Collaborate with Backend, QA, and Security teams to standardize release, rollback, and versioning processes; report infrastructure status, incidents, and cost to the CTO/Management regularly.
- Your Skills and Experience
Requirements
- Minimum 7 years of experience in DevOps/SRE/Infrastructure, including experience leading a team; experience spanning both cloud and on-premise environments is strongly preferred.
- Proficient with core AWS services (EKS, EC2, S3, RDS, IAM, VPC, ALB/NLB, CloudWatch) and Terraform (required).
- Hands-on experience managing on-premise infrastructure: Servers, Network, Storage, Virtualization (VMware/Hyper-V), Backup/DR.
- Strong understanding of Kubernetes operations at scale (EKS and/or on-prem), Helm, and multi-tenant architecture.
- Experience with CI/CD (Jenkins/GitLab CI/ArgoCD) and implementing observability (Prometheus/Grafana/ELK); experience defining SLIs/SLOs/SLAs.
- Solid understanding of system security (IAM/RBAC, secret management, least privilege, network/firewall) and FinOps/cost optimization.
- Experience operating database clusters, caching, message queues, search, or object storage (PostgreSQL/MySQL, Redis, Kafka/RabbitMQ, Elasticsearch, Ceph/MinIO) is a plus.
- Proven ability to lead/mentor a technical team, with strong communication skills across technical teams, senior stakeholders, and vendors.
- Why You'll Love Working Here
Benefits
- Competitive salary package based on qualifications and experience.
BonusCompany tripsTrainingLaptop / equipment
Summary of facts from the official posting. View original ↗
Interested in this role?
You'll be taken to the employer's official application page.
Apply on official site ↗
Is this your business?
Claim this page, request edits or removal
→