SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 200,000 - 400,000 / annual
Inferact, founded by the creators and core maintainers of vLLM, is seeking a hands-on cluster administration engineer to own and operate high-performance GPU compute infrastructure. The role focuses on ensuring cluster health, GPU availability, monitoring, alerting, scheduling, access control, diagnostics, and incident response across neo-cloud and dedicated compute providers.
Key responsibilities include:
- Taking ownership of cluster health and GPU availability across multiple infrastructure providers
- Implementing monitoring, alerting, and diagnostic systems for production GPU clusters
- Managing cluster scheduling and resource allocation using tools like SLURM or Kubernetes
- Owning urgent infrastructure incidents end-to-end when compute issues block engineering teams
- Automating operational workflows using Bash, Python, Ansible, Terraform, or Helm
- Working closely with engineering leadership to standardize provisioning, operations, and scaling across providers
- Debugging and resolving GPU-specific issues including driver management, memory errors, and hardware diagnostics
Required qualifications:
- Bachelor's degree or equivalent in computer science, engineering, systems administration, or related field
- Hands-on experience administering large compute clusters, HPC environments, university clusters, supercomputing systems, or production GPU clusters
- Strong Linux systems administration fundamentals (networking, processes, storage, package management, shell scripting, logs, access control, debugging)
- Experience operating GPU servers with knowledge of driver management, GPU health monitoring, node failures, and hardware diagnostics
- Experience with cluster scheduling tools (SLURM, Kubernetes, or equivalent)
- Ability to own urgent infrastructure incidents end-to-end
- Automation skills using Bash, Python, Ansible, Terraform, Helm, or similar tools
Preferred experience includes operating GPU compute across providers (Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod), improving cluster utilization, high-performance GPU networking (InfiniBand, RoCE, NVLink, RDMA, NCCL), HPC/ML storage systems (NFS, Lustre, Ceph), infrastructure security and access management, and background in research computing, ML infrastructure, SRE, or platform engineering.
Bonus qualifications include managing GPU/HPC infrastructure in university labs, national labs, research institutions, or AI infrastructure companies; building monitoring and remediation workflows; operating Kubernetes at scale for ML workloads; and standardizing operations across multiple providers.