SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff, Cluster Administration

Inferact - San Francisco, CA, USA - Hybrid - posted 2026-08-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 200,000 - 400,000 / annual

Inferact, founded by the creators and core maintainers of vLLM, is seeking a hands-on cluster administration engineer to own and operate high-performance GPU compute infrastructure. The role focuses on ensuring cluster health, GPU availability, monitoring, alerting, scheduling, access control, diagnostics, and incident response across neo-cloud and dedicated compute providers. Key responsibilities include: - Taking ownership of cluster health and GPU availability across multiple infrastructure providers - Implementing monitoring, alerting, and diagnostic systems for production GPU clusters - Managing cluster scheduling and resource allocation using tools like SLURM or Kubernetes - Owning urgent infrastructure incidents end-to-end when compute issues block engineering teams - Automating operational workflows using Bash, Python, Ansible, Terraform, or Helm - Working closely with engineering leadership to standardize provisioning, operations, and scaling across providers - Debugging and resolving GPU-specific issues including driver management, memory errors, and hardware diagnostics Required qualifications: - Bachelor's degree or equivalent in computer science, engineering, systems administration, or related field - Hands-on experience administering large compute clusters, HPC environments, university clusters, supercomputing systems, or production GPU clusters - Strong Linux systems administration fundamentals (networking, processes, storage, package management, shell scripting, logs, access control, debugging) - Experience operating GPU servers with knowledge of driver management, GPU health monitoring, node failures, and hardware diagnostics - Experience with cluster scheduling tools (SLURM, Kubernetes, or equivalent) - Ability to own urgent infrastructure incidents end-to-end - Automation skills using Bash, Python, Ansible, Terraform, Helm, or similar tools Preferred experience includes operating GPU compute across providers (Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod), improving cluster utilization, high-performance GPU networking (InfiniBand, RoCE, NVLink, RDMA, NCCL), HPC/ML storage systems (NFS, Lustre, Ceph), infrastructure security and access management, and background in research computing, ML infrastructure, SRE, or platform engineering. Bonus qualifications include managing GPU/HPC infrastructure in university labs, national labs, research institutions, or AI infrastructure companies; building monitoring and remediation workflows; operating Kubernetes at scale for ML workloads; and standardizing operations across multiple providers.

Similar roles