← Công ty TNHH Sapawoo
AI Infrastructure Engineer GPU & Colocation
Công ty TNHH Sapawoo · Hồ Chí Minh
Apply on official site ↗
Type
Full-time
Work mode
On-site
Level
Staff
Industry
Engineering / Mechanical
Salary
Thương lượng
Location
Hồ Chí Minh, Hồ Chí Minh, Hồ Chí Minh
Overview
- Plan and Build the Infrastructure
- Translate AI workload requirements into GPU, compute, storage, and networking configurations.
- Evaluate hardware compatibility, supplier proposals, and costs.
- Coordinate procurement, delivery, installation, and commissioning.
- Confirm rack space, power, cooling, connectivity, and access requirements with the colocation provider.
- Install and configure servers, networking equipment, cabling, and remote management.
- Maintain infrastructure documentation and asset inventories.
- Configure GPU and AI Systems
- Configure and maintain Linux servers, GPU drivers, CUDA environments, and container runtimes.
- Deploy and maintain model-serving infrastructure in collaboration with the AI and engineering teams.
- Diagnose GPU, memory, storage, and network performance issues.
- Benchmark workloads and improve resource utilisation.
- Automate deployment and configuration to make environments reproducible.
- Manage Networking and Security
- Configure network segmentation, firewalls, VPNs, and secure administrative access.
- Implement access controls, credential management, system hardening, and patching.
- Separate workloads and environments according to product and security requirements.
- Troubleshoot connectivity issues with infrastructure providers.
- Own Reliability and Operations
- Establish monitoring and alerting for hardware health, GPU utilisation, capacity, and service availability.
- Implement backup and recovery procedures and test restoration.
- Maintain operational runbooks and troubleshoot infrastructure incidents.
- Coordinate hardware replacements, maintenance windows, and remote-hands support.
- Identify single points of failure and propose practical improvements.
- Establish clear support coverage and escalation procedures with the team.
- Manage Capacity and Costs
- Track utilisation, operating costs, and capacity constraints.
- Recommend upgrades based on measured demand and performance.
- Assess trade-offs between owned hardware, rented GPU capacity, and cloud services.
- Plan expansion while avoiding unnecessary complexity and overprovisioning.
Summary of facts from the official posting. View original ↗
Interested in this role?
You'll be taken to the employer's official application page.
Apply on official site ↗
Is this your business?
Claim this page, request edits or removal
→