SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lambda is seeking an experienced HPC Support Engineer to serve as a senior technical escalation point for its AI cloud infrastructure platform. In this role, you'll troubleshoot complex infrastructure and platform issues at the hardware, driver, and kernel levels, working with tens of thousands of customers ranging from AI researchers to enterprises and hyperscalers.
Key responsibilities include: diagnosing root causes across distributed systems, GPU clusters, and HPC environments to distinguish between hardware failures, driver issues, kernel problems, and customer misconfiguration; proactively identifying and closing operational gaps in processes, tooling, and documentation; using AI tools to build scripts and automations that improve support efficiency; performing root-cause analysis on complex distributed systems; documenting solutions and evolving support procedures; collaborating with engineering teams to convert recurring customer pain points into permanent platform fixes; mentoring and training peer support engineers; and participating in a rotating on-call schedule to own major incidents and customer issues.
You'll need 3+ years of hands-on HPC experience in administration, support, or engineering roles, with very strong Linux system administration expertise. Required skills include proven experience in HPC environments with Linux cluster administration (strong preference for Kubernetes and/or Slurm), strong coding ability and CI/CD experience, proficiency with monitoring/logging tools (Prometheus, Grafana, Datadog), expertise in log analysis and kernel-level debugging, experience with CUDA, NCCL, NVLink, and GPUDirect RDMA, knowledge of high-throughput networking (IB/RoCE), understanding of distributed AI/ML and HPC workloads, and TCP/IP, VPN, and firewall knowledge in cloud environments. You should be able to work independently while mentoring junior engineers.
Nice-to-have qualifications include experience with virtualization and container technologies (Docker, Kubernetes), prior work with GPU cloud providers, flexible availability for off-hours shifts, high-performance storage systems knowledge, infrastructure-as-code tools (Terraform, Ansible), and NVIDIA GPU and Infiniband experience.