← Công ty TNHH Sapawoo
AI Infrastructure Engineer GPU & Colocation
Công ty TNHH Sapawoo · Hồ Chí Minh
Ứng tuyển tại trang chính thức ↗
Loại hình
Toàn thời gian
Hình thức
Tại văn phòng
Cấp bậc
Nhân viên
Ngành nghề
Kỹ thuật / Cơ khí
Mức lương
Thương lượng
Địa điểm
Hồ Chí Minh, Hồ Chí Minh, Hồ Chí Minh
Tổng quan
- Plan and Build the Infrastructure
- Translate AI workload requirements into GPU, compute, storage, and networking configurations.
- Evaluate hardware compatibility, supplier proposals, and costs.
- Coordinate procurement, delivery, installation, and commissioning.
- Confirm rack space, power, cooling, connectivity, and access requirements with the colocation provider.
- Install and configure servers, networking equipment, cabling, and remote management.
- Maintain infrastructure documentation and asset inventories.
- Configure GPU and AI Systems
- Configure and maintain Linux servers, GPU drivers, CUDA environments, and container runtimes.
- Deploy and maintain model-serving infrastructure in collaboration with the AI and engineering teams.
- Diagnose GPU, memory, storage, and network performance issues.
- Benchmark workloads and improve resource utilisation.
- Automate deployment and configuration to make environments reproducible.
- Manage Networking and Security
- Configure network segmentation, firewalls, VPNs, and secure administrative access.
- Implement access controls, credential management, system hardening, and patching.
- Separate workloads and environments according to product and security requirements.
- Troubleshoot connectivity issues with infrastructure providers.
- Own Reliability and Operations
- Establish monitoring and alerting for hardware health, GPU utilisation, capacity, and service availability.
- Implement backup and recovery procedures and test restoration.
- Maintain operational runbooks and troubleshoot infrastructure incidents.
- Coordinate hardware replacements, maintenance windows, and remote-hands support.
- Identify single points of failure and propose practical improvements.
- Establish clear support coverage and escalation procedures with the team.
- Manage Capacity and Costs
- Track utilisation, operating costs, and capacity constraints.
- Recommend upgrades based on measured demand and performance.
- Assess trade-offs between owned hardware, rented GPU capacity, and cloud services.
- Plan expansion while avoiding unnecessary complexity and overprovisioning.
Tóm tắt thông tin từ tin tuyển dụng chính thức. Xem bản gốc ↗
Quan tâm đến vị trí này?
Bạn sẽ được chuyển đến trang ứng tuyển chính thức của nhà tuyển dụng.
Ứng tuyển tại trang chính thức ↗
Đây là doanh nghiệp của bạn?
Nhận quản lý trang, yêu cầu chỉnh sửa hoặc gỡ bỏ
→