← NAB Innovation Centre Vietnam
Senior/Middle Site Reliability Engineer
NAB Innovation Centre Vietnam · Hồ Chí Minh
Apply on official site ↗
Type
Full-time
Work mode
On-site
Level
Staff
Industry
Engineering / Mechanical
Salary
Thương lượng
Location
Thành phố Thủ Đức, Hồ Chí Minh, Hồ Chí Minh
Overview
- Provide support to customers of our services during business hours and in a rotating support roster due to the 24x7 nature of the services
- Work on Change Requests, Incidents, Defects and Problem Records
- Managed the full lifecycle of application deployment, successfully executing releases and hotfixes per year using automated CI/CD pipelines built with Jenkins and Harness.
- Acted as the lead responder for 24x7 support, resolving production incidents annually and driving post-mortem analysis to strengthen system resilience.
- Served as a key technical point of contact for banking staff, investigating and resolving business-critical issues to ensure operational continuity for internal stakeholders.
- Engineered a seamless migration from AWS ECS to EKS, enhancing the platform's scalability and maintainability with zero disruption to banking services.
- Implement and maintain monitoring solutions for Web applications and microservices architecture.
- Built and scaled a sophisticated observability stack using Splunk, CloudWatch, and AppDynamics, and led the successful migration of critical alerting systems from Splunk to the OpenSearch platform to enhance query performance and reduce operational costs.
- Fulfil other tasks as assigned by your People Leader and/or authorized representative of NAB Vietnam from time to time.
- Your Skills and Experience
- 4+ years strong experience in SRE & DevOps/Platform domain with weight 80% SRE + 20% DevOps/Platform knowledge
- Experience in investigating & tracing issue, problem, incident, and customer support
- Knowledge in Incident lifecycle management; Runbook & automation; Communication & Stakeholder management; Trouble shooting & system thinking, etc
- Hands-on experience with monitoring tools such as AppD, OpenSearch, AWS dashboards, and alerts to maintain service stability.
- Able to collaborate with technical teams, third-party vendors, upstream/downstream team to uphold service quality from a customer-centric perspective and ensure seamless service integration
- Able to mentor other members & build up practice for team
- Good English skills, both oral and written.
- Experienced in AWS or Azure cloud technologies
- Experience in building CI/CD pipeline automation, tooling (Github, Jenkins, Artifactory, and Docker) and Compliance as code
- Nice to have:
- Familiarity with ITIL processes and on-call operations, troubleshooting on production environment
- Hands-on Kubernetes experience and strong knowledge of core AWS services (EC2, ECS, EKS, S3, DynamoDB, Lambda, CloudWatch, SQS, SNS).
- Why You'll Love Working Here
Benefits
- 20-day annual leave and 7-day sick leave, etc.
BonusHealthcareStock options / ESOP
Summary of facts from the official posting. View original ↗
Interested in this role?
You'll be taken to the employer's official application page.
Apply on official site ↗
Is this your business?
Claim this page, request edits or removal
→