Skip to content
Tuyển Dụng
← Pizza Hut Digital & Technology

Site Reliability Engineer SRE, AWS, Azure, Cloud

Pizza Hut Digital & Technology · Hồ Chí Minh
Apply on official site ↗
Type
Full-time
Work mode
On-site
Level
Staff
Industry
Engineering / Mechanical
Salary
Thương lượng
Location
Quận Tân Bình, Hồ Chí Minh, Hồ Chí Minh

Overview

  • Independently lead complex incidents involving multiple systems, teams, or dependencies.
  • Coordinate incident response activities, facilitate communication, and drive timely resolution.
  • Lead or contribute to post-incident reviews and root cause analysis activities.
  • Ensure corrective and preventive actions are identified, prioritized, tracked, and completed
  • Design, implement, and continuously optimize monitoring, logging, alerting, and tracing solutions.
  • Develop meaningful alerts based on service behavior, customer impact, and business priorities.
  • Build and maintain dashboards that provide actionable insights into system performance and reliability.
  • Own SRE responsibilities for one or more markets, platforms, or critical services end-to-end.
  • Establish and maintain operational excellence standards for assigned domains.
  • Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date.
  • Continuously assess platform health, identify reliability risks, and drive improvements before incidents occur.
  • Partner with engineering teams to ensure new features and services meet reliability requirements before production release.
  • Define and track reliability metrics, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Develop and maintain tools, scripts, and automation solutions that reduce manual effort and improve operational efficiency.
  • Identify and eliminate repetitive tasks through automation and self-service capabilities.
  • Establish and promote best practices for the responsible use of AI within SRE workflows
  • Mentor and support Level 5 and Level 6 engineers in incident management, monitoring, automation, AI adoption, and operational best practices.
  • Review monitoring configurations, dashboards, runbooks, and operational documentation to maintain quality standards.
  • Share knowledge through training sessions, documentation, and post-incident learning activities.
  • Contribute to the continuous improvement of team processes, standards, and ways of working.
  • Develop meaningful alerts based on service behaviours, customer impact, and business priorities.
  • Influence technical decisions that improve platform stability, scalability, and operational efficiency

Benefits

BonusTraining

Summary of facts from the official posting. View original ↗

Interested in this role?

You'll be taken to the employer's official application page.

Apply on official site ↗
Is this your business? Claim this page, request edits or removal