Đến nội dung chính
Tuyển Dụng
← CÔNG TY TRÁCH NHIỆM HỮU HẠN NANYANG BIOLOGICS VIỆT NAM

Senior MLOps Engineer (GPU Platform)

CÔNG TY TRÁCH NHIỆM HỮU HẠN NANYANG BIOLOGICS VIỆT NAM · Hà Nội
Ứng tuyển tại trang chính thức ↗
Loại hình
Toàn thời gian
Hình thức
Tại văn phòng
Cấp bậc
Nhân viên
Ngành nghề
Kỹ thuật / Cơ khí
Mức lương
Đến 179tr
Địa điểm
Thành phố Hà Nội, Hà Nội

Tổng quan

  • Own the success-run ratio of GPU workloads as a measurable SLO; drive it up and keep it there.
  • Build and operate the GPU job scheduling and queueing layer — fair-share allocation, prioritization, backpressure, and recovery across a heterogeneous fleet.
  • Implement GPU partitioning and sharing (MIG, MPS, time-slicing) to raise utilization without destabilizing runs.
  • Profile and right-size workloads: per-model GPU memory, runtime, and failure characteristics; eliminate OOMs and silent failures.
  • Define a standard packaging/deployment contract for new models so onboarding is repeatable, not bespoke.
  • Build observability for the run lifecycle — metrics, logs, traces, alerting — so failures are caught and diagnosed fast.
  • Harden the orchestration stack (workflow engine, durable execution, retries/failover) against real failure modes.
  • Partner with the DevOps engineer on cluster/networking and with AI engineers to make their models production-ready.
  • 5+ years in MLOps / ML platform / GPU systems engineering, with direct ownership of production reliability.
  • Deep experience operating GPU workloads at scale (NVIDIA stack: CUDA, drivers, GPU Operator, MIG/MPS).
  • Strong background in workload orchestration and scheduling — Kubernetes (Jobs/batch), Ray, Slurm, or equivalent.
  • Hands-on managed-ML platform experience on at least one major cloud, with working familiarity of the other:
  • GCP — Cloud Run, Vertex AI
  • AWS — SageMaker
  • Solid understanding of cloud architecture (compute, networking, storage, IAM) across hybrid cloud + on-prem.
  • Proven track record raising reliability/utilization of a heterogeneous GPU fleet.
  • Solid software engineering (Python and one systems language) — you build platform tooling, not just configure it.
  • Observability and SRE fundamentals: SLOs, metrics, tracing, incident response.

Tóm tắt thông tin từ tin tuyển dụng chính thức. Xem bản gốc ↗

Quan tâm đến vị trí này?

Bạn sẽ được chuyển đến trang ứng tuyển chính thức của nhà tuyển dụng.

Ứng tuyển tại trang chính thức ↗
Đây là doanh nghiệp của bạn? Nhận quản lý trang, yêu cầu chỉnh sửa hoặc gỡ bỏ