← Pizza Hut Digital & Technology
Site Reliability Engineer SRE, AWS, Azure, Cloud
Pizza Hut Digital & Technology · Hồ Chí Minh
Apply on official site ↗
Type
Full-time
Work mode
On-site
Level
Staff
Industry
Engineering / Mechanical
Salary
Thương lượng
Location
Quận Tân Bình, Hồ Chí Minh, Hồ Chí Minh
Overview
- Independently lead complex incidents involving multiple systems, teams, or dependencies.
- Coordinate incident response activities, facilitate communication, and drive timely resolution.
- Lead or contribute to post-incident reviews and root cause analysis activities.
- Ensure corrective and preventive actions are identified, prioritized, tracked, and completed
- Design, implement, and continuously optimize monitoring, logging, alerting, and tracing solutions.
- Develop meaningful alerts based on service behavior, customer impact, and business priorities.
- Build and maintain dashboards that provide actionable insights into system performance and reliability.
- Own SRE responsibilities for one or more markets, platforms, or critical services end-to-end.
- Establish and maintain operational excellence standards for assigned domains.
- Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date.
- Continuously assess platform health, identify reliability risks, and drive improvements before incidents occur.
- Partner with engineering teams to ensure new features and services meet reliability requirements before production release.
- Define and track reliability metrics, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Develop and maintain tools, scripts, and automation solutions that reduce manual effort and improve operational efficiency.
- Identify and eliminate repetitive tasks through automation and self-service capabilities.
- Establish and promote best practices for the responsible use of AI within SRE workflows
- Mentor and support Level 5 and Level 6 engineers in incident management, monitoring, automation, AI adoption, and operational best practices.
- Review monitoring configurations, dashboards, runbooks, and operational documentation to maintain quality standards.
- Share knowledge through training sessions, documentation, and post-incident learning activities.
- Contribute to the continuous improvement of team processes, standards, and ways of working.
- Develop meaningful alerts based on service behaviours, customer impact, and business priorities.
- Influence technical decisions that improve platform stability, scalability, and operational efficiency
Benefits
BonusTraining
Summary of facts from the official posting. View original ↗
Interested in this role?
You'll be taken to the employer's official application page.
Apply on official site ↗
Is this your business?
Claim this page, request edits or removal
→