Genesis Networks Pte Ltd · MyCareersFuture · 1mo
AI Infrastructure Engineer
Singapore, Central- Posted
- 2026-08-14 (1mo)
- Place
- Singapore, Central
- Commitment
- Full Time
- Salary
- SGD 6,400 – 7,000 / month
- Experience
- 3+ YOE
- Education
- Bachelor's
- Department
- Information Technology
- Source
- MyCareersFuture (the employer’s own listing)
Your match
Sign in to see how your skills match this job.
Skills in this posting
PythonShell ScriptingBashAmazon Web ServicesGoogle Cloud PlatformKubernetesHelmTerraformAnsibleInfrastructure as CodeIaCPrometheus
1. Compute & ClusterManagement
• Architect, configure, and maintain high-density multi-GPU compute clusters (e.g., NVIDIA HGX/DGX architectures).
• Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.
• Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.
2. High-Performance Networking& Storage
• Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).
• Configure and scale high-throughput parallel file systems and object storage (e.g., Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed data pipelines.
3. Automation &Infrastructure as Code (IaC)
• Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi .
• Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.
4. Operations, Observability& Performance
• Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).
• Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.
• Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.
Qualifications &Requirements
Technical Competencies
• Operating Systems: Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
• Accelerated Compute: Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
• Orchestration & Workload Scheduling: Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).
• High-Speed Networking: Proven experience with RDMA (RoCE v2 / InfiniBand), PFC (Priority Flow Control), and ECN configurations.
• Storage Systems: Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.
• Automation: Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).
Experience & Education
• Bachelor’s Degree in Computer Science, Information Technology, Computer Engineering, or equivalent practical experience.
• 3–6+ years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.
• Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified Associate/Professional, AWS/Azure/GCP Solutions Architect).
Also posted at mycareersfuture.gov.sg, mycareersfuture.gov.sg
Genesis Networks Pte Ltd
6 open roles in Singapore, straight from Genesis Networks Pte Ltd’s own careers page.
- AI Application DeveloperSingapore, Central · 1mo
- Head of Security Operations & Delivery (IT)Singapore, Central · 1mo
- AI Security & Automation LeadSingapore, Central · 1mo
- Sales Manager (IT)Singapore, Central · 2mo
- ICT Security Engineer / Service Desk SpecialistSingapore, Central · 2mo