Position Summary
We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization.
The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.
Responsibilities
-
Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.
-
Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.
-
Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.
-
Build and support production AI infrastructure running hundreds to thousands of GPUs.
-
Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.
-
Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.
-
Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.
-
Configure and tune distributed AI software stacks including:
- PyTorch
- NCCL
- CUDA
- UCX
- MPI
- Slurm
- Pyxis/Enroot
-
Optimize GPU scheduling and resource allocation for both training and inference environments.
-
Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.
-
Identify performance regressions and troubleshoot distributed training issues at scale.
-
Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.
-
Work closely with ML engineers to improve training scalability and inference efficiency.
-
Create automation to deploy, validate, benchmark, and monitor GPU clusters.
-
Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.
Required Qualifications
-
7+ years designing or operating large-scale Linux infrastructure.
-
5+ years supporting production GPU clusters for AI or HPC workloads.
-
Demonstrated experience building multi-node GPU training environments from the ground up.
-
Deep expertise with distributed PyTorch training.
-
Extensive experience troubleshooting and optimizing NCCL communications.
-
Strong understanding of distributed AI communication patterns, including:
- AllReduce
- ReduceScatter
- AllGather
- Broadcast
- Point-to-point communications
-
Experience benchmarking distributed training using tools such as:
- nccl-tests
- NVIDIA DCGM
- Nsight Systems
- MLPerf (preferred)
-
Strong understanding of GPU memory management, including:
- KV Cache
- Activation checkpointing
- Tensor Parallelism
- Pipeline Parallelism
- Data Parallelism
-
Experience optimizing LLM inference throughput, including:
- Tokens/sec optimization
- Batch sizing
- Continuous batching
- KV cache tuning
- Memory bandwidth optimization
-
Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.
-
Expert-level Linux systems administration skills.
-
Experience with Slurm workload manager.
-
Experience using Pyxis and Enroot for containerized GPU workloads.
-
Strong scripting skills using Python and Bash.
Technical Expertise
AI Frameworks
- PyTorch
- CUDA
- NCCL
- Triton (preferred)
- TensorRT-LLM (preferred)
Cluster Scheduling
- Slurm
- Pyxis
- Enroot
GPU Networking
Strong understanding of:
- InfiniBand
- RoCE v2
- RDMA
- GPUDirect RDMA
- GPUDirect Storage
- UCX
- MPI
- Network topology optimization
- Congestion control
- QoS
- ECN/PFC
- High-speed Ethernet (200/400/800 GbE)
Storage
Experience designing or tuning storage for AI workloads, including:
- Parallel file systems
- Distributed storage
- Object storage
- NVMe
- Checkpoint optimization
- Dataset staging
- GPUDirect Storage
- Storage bandwidth optimization
- Metadata performance
Performance Engineering
Experience with:
- NCCL benchmarking
- Multi-node scaling analysis
- GPU utilization optimization
- Communication/computation overlap
- NUMA optimization
- CPU affinity
- PCIe topology
- GPU topology (NVLink/NVSwitch)
- Memory bandwidth analysis
- End-to-end performance profiling
Preferred Qualifications
- Experience deploying AI workloads on Kubernetes.
- Experience with NVIDIA GPU Operator.
- Experience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.).
- Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.
- Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.
- Familiarity with MLPerf benchmarking.
- Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.
- Experience automating infrastructure using Ansible, Terraform, or similar tools.
- Experience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.