← Công ty Cổ phần Thanh toán số MobiFone
Senior Manager Service Reliability Engineer 24/7 Team
Công ty Cổ phần Thanh toán số MobiFone · Hà Nội
Ứng tuyển tại trang chính thức ↗
Loại hình
Toàn thời gian
Hình thức
Tại văn phòng
Cấp bậc
Trưởng / Phó phòng
Ngành nghề
Kỹ thuật / Cơ khí
Mức lương
Thương lượng
Địa điểm
Quận Ba Đình, Hà Nội, Hà Nội
Tổng quan
- Incident Management & Command
- Act as Incident Commander with end-to-end accountability for high-severity production incidents (P1/P2).
- Lead the war-room and coordinate cross-functional teams (Application, Infrastructure, Network, Security, Vendor) throughout incident resolution.
- Ensure the fastest possible service restoration, minimizing MTTR (Mean Time To Recovery) and reducing impact to customers and business operations.
- Maintain clear, timely communication with stakeholders (business, leadership, operations) across the full incident lifecycle.
- Conduct Post-Incident Reviews (PIR), Root Cause Analysis (RCA) and define preventive actions.
- Service Operations (24/7)
- Operate mission-critical / high-availability systems on a continuous 24/7 basis.
- Establish, monitor and enforce SLA / SLO and compliance across critical services.
- Manage the Incident / Problem / Change lifecycle following the ITIL/ITSM framework.
- Improve system stability, reliability and resilience.
- 24/7 Operations & Shift Model
- Operate under a 24/7 model organized as shifts/teams, ensuring round-the-clock coverage including nights, weekends and public holidays.
- Plan, manage and balance shift rosters and on-call schedules to guarantee adequate staffing and seamless handover between shifts.
- Lead or participate in the duty-officer / on-call rotation as Incident Commander, ready to respond to major incidents at any time.
- Ensure each shift maintains complete handover logs, runbooks and shift reports for full operational continuity.
- Willing and able to work in rotating shifts (day/night) and respond outside office hours when major incidents occur.
- Monitoring & Observability
- Design and operate comprehensive monitoring systems (APM, infrastructure monitoring, log/observability).
- Proactively detect issues through alert tuning and early warning signals.
- Optimize dashboards and alerting strategy to reduce noise and false alerts.
- Problem Management & Continuous Improvement
- Analyze incident trends and identify recurring issues.
- Implement automation and runbooks to reduce manual effort and resolution time.
- Propose architecture / system design improvements to increase availability.
- Leadership & Stakeholder Management
- Manage, lead and mentor the IT Operations / Incident Management team, including shift supervisors and on-call engineers.
- Act as the bridge between Business / Product / Engineering / Infrastructure.
- Report directly to senior leadership on overall system health and status.
- Your Skills and Experience
- 7–12+ years of experience in IT Operations / ITSM / Incident Management.
- Hands-on experience handling major incidents in a 24/7 operating environment.
- Background in Banking / Fintech / Payment Gateway / E-commerce or other large-scale systems.
- Strong understanding of ITIL / ITSM best practices.
- Willing to work in a 24/7 shift-based model (3 shifts / 4 teams) and participate in on-call duty rotations.
- Solid infrastructure foundation: Network (TCP/IP, Load Balancing, DNS), Server (Linux/Windows), basic Storage / Database.
- Proficient with monitoring tools such as Datadog, Dynatrace, Prometheus, Grafana, ELK, Splunk, etc.
- Experience with distributed systems, microservices, high availability / DR / failover.
- Nice-to-have
- Experience building an Incident Management framework.
Quyền lợi
Chế độ thưởngDu lịchLaptop / Thiết bị
Tóm tắt thông tin từ tin tuyển dụng chính thức. Xem bản gốc ↗
Quan tâm đến vị trí này?
Bạn sẽ được chuyển đến trang ứng tuyển chính thức của nhà tuyển dụng.
Ứng tuyển tại trang chính thức ↗
Đây là doanh nghiệp của bạn?
Nhận quản lý trang, yêu cầu chỉnh sửa hoặc gỡ bỏ
→