Đến nội dung chính
Tuyển Dụng
← Techcombank

Expert, Site Reliability Engineer (Urgent) - Techcomlife

Techcombank · Hà Nội
Ứng tuyển tại trang chính thức ↗
Loại hình
Toàn thời gian
Hình thức
Tại văn phòng
Cấp bậc
Nhân viên
Ngành nghề
Kỹ thuật / Cơ khí
Mức lương
Thương lượng
Địa điểm
Hà Nội, Hà Nội, Hà Nội

Tổng quan

  • Perform specialized tasks, provide special skills in daily monitoring of IT infrastructure / applications / services of critical services (ROC, COC,...) to ensure that critical services meet SLAs committed to the business; As well as in the process of handling alerts, try to restore services as quickly as possible as well as remaining issues to ensure the best service delivery to customers.
  • KEY ACCOUNTABILITIES
  • Key Accountabilities (1)
  • Participate in monitoring and handling system alerts/incidents/problems:
  • Perform 24/7 monitoring and handle alerts of services of the entire IT infrastructure/application/services. In case encounter difficulties, escalate to L3 for coordinated processing.
  • Ensure projects/specialized operations departments provide adequate alert/incident handling instructions for new services before golive and periodically review and update existing alert/incident handling instructions.
  • Responsible for periodically reviewing issues/vulnerabilities in IT infrastructure/applications/services within scope of responsibility
  • Provide in-depth transfer skills in monitoring and handling alerts and critical IT service incidents
  • Participate Lead the standardizing and developing relevant processes and regulations to ensure effective monitoring and handling of alerts/incidents.
  • Coordinate with relevant units to promptly restore services/systems, investigate root causes, propose solutions and implement solutions.
  • Participate in implementing changes across the software development environment, including on Prem and cloud.
  • Participate in building and optimizing centralized monitoring tools:
  • Implement the development and promulgation of standards and operate centralized monitoring tools (Dynatrace, Grafana, Splunk...)
  • Implement monitoring tool integration and support building monitoring dashboards for new IT infrastructure/applications/services
  • Ensure projects/specialized operations departments provide adequate monitoring indicators/monitoring thresholds for new services before golive.
  • Key Accountabilities (2)
  • System problem and incident management:
  • Manage the lifecycle of IT incidents, including identifying, classifying, coordinating and resolving incidents according to SLAs
  • Be the contact point during troubleshooting, ensuring effective communication among technical, operations and business departments
  • Root cause analysis (RCA) after each incident, recommending preventive measures and process improvements. Coordinate with relevant teams to minimize downtime and improve system availability.
  • Participate in developing and maintaining incident management processes according to standards and best practices
  • Key Accountabilities (3)
  • Control and ensure the unit's activities comply with issued policies, regulations, procedures and instructions.
  • Identify the unit's risks during operations, coordinate with relevant units to develop methods to measure, evaluate and minimize risks.
  • Report periodically to management levels and perform other tasks as directed by management

Tóm tắt thông tin từ tin tuyển dụng chính thức. Xem bản gốc ↗

Quan tâm đến vị trí này?

Bạn sẽ được chuyển đến trang ứng tuyển chính thức của nhà tuyển dụng.

Ứng tuyển tại trang chính thức ↗
Đây là doanh nghiệp của bạn? Nhận quản lý trang, yêu cầu chỉnh sửa hoặc gỡ bỏ