SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lambda is an AI cloud infrastructure company serving tens of thousands of customers from AI researchers to enterprises and hyperscalers. The Senior Site Reliability Engineer for Fleet role focuses on building and operating large-scale HPC clusters optimized for AI workloads.
Key responsibilities include designing and maintaining monitoring and alerting systems for cluster health across fabric, GPU, power/thermal, and job-level signals. You'll remotely deploy and configure large-scale HPC clusters using automation tools like Ansible and Terraform, treating infrastructure as code. The role involves automating the full cluster lifecycle including operating systems, firmware, drivers, and networking configurations.
You'll create runbooks and automated remediations for common cluster failure modes, enabling Support teams to respond safely and efficiently. Troubleshooting spans InfiniBand/RoCE, NCCL, GPU-direct, fabric switching, and power systems, working closely with on-site deployment teams. The position includes on-call rotations and leading incident response for cluster-level issues.
Required qualifications include 7+ years in Site Reliability Engineering, HPC Engineering, DevOps, or similar roles. You need strong understanding of modern AI infrastructure from GPU architectures to hardware performance optimization, plus deep Linux expertise in distributed environments. Experience configuring and troubleshooting InfiniBand, RoCE, CLOS fabrics, 100GbE Ethernet, GPU-direct, and NCCL is essential. Solid Python and Go skills are required, along with proficiency in monitoring tools (Prometheus, Grafana, Clickhouse) and infrastructure automation (Ansible, Terraform).
Nice-to-have skills include experience with ML frameworks (PyTorch, TensorFlow), containerization (Docker, Kubernetes), HPC operations, NVIDIA hardware/firmware depth, data center power/thermal design, chaos engineering, and compliance frameworks (SOC 2, ISO 27001).
Lambda was founded in 2012 and has 500+ employees. The company is backed by notable investors including NVIDIA, Andrej Karpathy, ARK Invest, In-Q-Tel, and others. The position requires 4 days per week in the San Francisco office (Fremont St location), with Tuesday designated as work-from-home day.