SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lambda is building AI cloud infrastructure for the next generation of AI training and inference at scale. This Senior Software Engineer role focuses on the Managed Kubernetes (Mk8s) team, responsible for designing and operating a purpose-built Kubernetes platform optimized for GPU-accelerated AI workloads—comparable to GKE but tailored for bare-metal AI infrastructure.
You will design, build, and maintain scalable control plane services, custom Kubernetes controllers, and operators. Core responsibilities include developing end-to-end cluster lifecycle management automation (provisioning, upgrades, patching, deletion) in Go and Python; building GPU-aware orchestration systems with intelligent GPU scheduling and resource allocation; and partnering with the Network team on high-performance networking solutions including CNI integration (Cilium, Multus), InfiniBand, RoCE, RDMA, and GPUDirect.
You'll develop resilient distributed systems that handle failure gracefully across large-scale environments, build platform services for inference (model serving, autoscaling, multi-model deployment), create internal tools and CLIs for ML/AI teams, and participate in on-call production support.
Required: 6+ years software engineering experience with significant technical scope ownership; deep Kubernetes internals knowledge (controllers, schedulers, operators, CRDs, CSI, CNI); distributed systems fundamentals; experience operating large-scale Kubernetes control planes; observability expertise (Prometheus, Grafana, distributed tracing); strong Go and Python programming; solid Linux, networking, containers, and cloud infrastructure knowledge.
Preferred: managed Kubernetes service experience (GKE, EKS, AKS); NVIDIA GPU/networking ecosystem hands-on work (GPU Operator, device plugins, DCGM, MIG, Network Operator, NCCL); HPC and job scheduler familiarity (Slurm, KAI, Volcano, Kueue); GPU/InfiniBand/RDMA/HPC experience; AI/ML storage architecture exposure; CNCF or Kubernetes SIG contributions.
Lambda is founded by ML practitioners and serves leading AI research labs and enterprises. The role offers exposure across the full infrastructure stack (Kubernetes, networking, storage, compute), deep NVIDIA partnership integration, and direct impact on systems powering global AI breakthroughs. Hybrid role requires 4 days/week in San Francisco, San Jose, or Bellevue office; Tuesday is designated work-from-home.