Đến nội dung chính
Tuyển Dụng
← Công ty TNHH Sapawoo

AI Infrastructure Engineer GPU & Colocation

Công ty TNHH Sapawoo · Hồ Chí Minh
Ứng tuyển tại trang chính thức ↗
Loại hình
Toàn thời gian
Hình thức
Tại văn phòng
Cấp bậc
Nhân viên
Ngành nghề
Kỹ thuật / Cơ khí
Mức lương
Thương lượng
Địa điểm
Hồ Chí Minh, Hồ Chí Minh, Hồ Chí Minh

Tổng quan

  • Plan and Build the Infrastructure
  • Translate AI workload requirements into GPU, compute, storage, and networking configurations.
  • Evaluate hardware compatibility, supplier proposals, and costs.
  • Coordinate procurement, delivery, installation, and commissioning.
  • Confirm rack space, power, cooling, connectivity, and access requirements with the colocation provider.
  • Install and configure servers, networking equipment, cabling, and remote management.
  • Maintain infrastructure documentation and asset inventories.
  • Configure GPU and AI Systems
  • Configure and maintain Linux servers, GPU drivers, CUDA environments, and container runtimes.
  • Deploy and maintain model-serving infrastructure in collaboration with the AI and engineering teams.
  • Diagnose GPU, memory, storage, and network performance issues.
  • Benchmark workloads and improve resource utilisation.
  • Automate deployment and configuration to make environments reproducible.
  • Manage Networking and Security
  • Configure network segmentation, firewalls, VPNs, and secure administrative access.
  • Implement access controls, credential management, system hardening, and patching.
  • Separate workloads and environments according to product and security requirements.
  • Troubleshoot connectivity issues with infrastructure providers.
  • Own Reliability and Operations
  • Establish monitoring and alerting for hardware health, GPU utilisation, capacity, and service availability.
  • Implement backup and recovery procedures and test restoration.
  • Maintain operational runbooks and troubleshoot infrastructure incidents.
  • Coordinate hardware replacements, maintenance windows, and remote-hands support.
  • Identify single points of failure and propose practical improvements.
  • Establish clear support coverage and escalation procedures with the team.
  • Manage Capacity and Costs
  • Track utilisation, operating costs, and capacity constraints.
  • Recommend upgrades based on measured demand and performance.
  • Assess trade-offs between owned hardware, rented GPU capacity, and cloud services.
  • Plan expansion while avoiding unnecessary complexity and overprovisioning.

Tóm tắt thông tin từ tin tuyển dụng chính thức. Xem bản gốc ↗

Quan tâm đến vị trí này?

Bạn sẽ được chuyển đến trang ứng tuyển chính thức của nhà tuyển dụng.

Ứng tuyển tại trang chính thức ↗
Đây là doanh nghiệp của bạn? Nhận quản lý trang, yêu cầu chỉnh sửa hoặc gỡ bỏ
→