← Molex Vietnam Co., Ltd.
Senior DevOps Engineer /SRE
Molex Vietnam Co., Ltd. · Hà Nội
Ứng tuyển tại trang chính thức ↗
Loại hình
Toàn thời gian
Hình thức
Tại văn phòng
Cấp bậc
Nhân viên
Ngành nghề
Kỹ thuật / Cơ khí
Mức lương
Thương lượng
Địa điểm
Hà Nội, Hà Nội, Hà Nội
Tổng quan
- We are hiring a Senior Observability Engineer / Platform Owner to drive enterprise observability enablement and ongoing operational visibility across critical digital manufacturing systems. This role will lead incoming dashboard, log, trace, alerting, and telemetry requests from ETS, MES, UFE, APS, AIP, SAP, Teamcenter, and related teams, turning them into practical, scalable, and maintainable…
- Serve as the single intake point for observability requests across ETS, MES, UFE, APS, AIP, SAP, Teamcenter, and related teams.
- Design and break down solutions covering dashboards, alerts, logs, traces, matrix configuration, and telemetry collection.
- Evaluate priority, delivery sequencing, dependencies, and effort, and maintain the backlog.
- Review deliverables from 2 interns to ensure quality, consistency, reuse, and maintainability.
- Handle complex topics such as cross-system tracing, log ingestion boundaries, metric definitions, and alert threshold design.
- Drive key initiatives such as HVR sync alerting, MES logs to Loki, Alloy standardization, and Teamcenter / SAP observability visibility.
- Build and maintain runbooks, handover documentation, templates, and configuration standards.
- Work closely with K8s / SRE, DBA, Windows host admins, and application owners to land observability solutions.
- Lead known issue management, top error log analysis, and rule design for alert-to-known-issue closure.
- Define use cases for distributed trace support dashboards, trace detail workflows, and AI-assisted analysis support.
- Govern the boundary between Grafana-managed alerts and data source-managed alerts, including notification paths, Alertmanager configuration ownership, and template strategy.
- Design solutions for complex dashboard data models such as QMS multi-table queries, panel reuse, hidden variables, and GUI embedding constraints.
- Define onboarding and integration approaches for external metrics such as Confluent Kafka metrics, AdminTools-based metrics, HVR sync metrics, and IIoT / Ignition health metrics.
- As a longer-term evolution target, help transition the observability platform toward AIOps capability by preparing usable data foundations for anomaly detection, AI-assisted RCA, and intelligent alerting.
- Drive the capture and organization of key AIOps inputs such as metrics, logs, alerts, application dependency trees / graphs, application change events, historical issues, and RCA records.
- Work with SRE, ETS, and application teams to establish feedback loops for auto-RCA outputs, operational knowledge capture, and future optimization.
- Identify which alerting, log analysis, and known issue scenarios are suitable for future anomaly detection, intelligent analysis, or agent-assisted troubleshooting, without disrupting current observability foundation priorities.
Tóm tắt thông tin từ tin tuyển dụng chính thức. Xem bản gốc ↗
Quan tâm đến vị trí này?
Bạn sẽ được chuyển đến trang ứng tuyển chính thức của nhà tuyển dụng.
Ứng tuyển tại trang chính thức ↗
Đây là doanh nghiệp của bạn?
Nhận quản lý trang, yêu cầu chỉnh sửa hoặc gỡ bỏ
→