SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lambda is building the AI Cloud infrastructure of the future. As a Senior Software Engineer on the Managed Kubernetes (Mk8s) team, you will design and build the orchestration platform that powers mission-critical AI workloads at scale—think GKE, but purpose-built for AI and running on bare metal.
You will work at the intersection of distributed systems, GPU-accelerated computing, and cloud-native infrastructure. Your responsibilities include:
- Design, build, and maintain scalable control plane services, operators, and custom Kubernetes controllers; develop automation in Go/Python for end-to-end cluster lifecycle management (provisioning, upgrades, patching, deletion)
- Build GPU-aware orchestration systems supporting GPU scheduling and resource allocation within the platform architecture
- Partner with the Network team on networking solutions for AI workloads: CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and GPUDirect
- Write resilient systems that handle failure gracefully—timeouts, retries, backoff, degraded-mode operation—across large-scale distributed environments
- Develop platform services for inference: model serving infrastructure, autoscaling based on inference load, multi-model deployment patterns
- Build internal tools and CLIs enabling ML/AI teams to deploy and monitor their own inference services
- Support and debug production issues through on-call rotation
Lambda serves tens of thousands of customers ranging from AI researchers to enterprises and hyperscalers. The company was founded in 2012, has 500+ employees, and is backed by notable investors including NVIDIA, Andrej Karpathy, ARK Invest, and In-Q-Tel. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.
This role requires presence in the San Francisco, San Jose, or Bellevue office 4 days per week; Tuesday is the designated work-from-home day.
REQUIREMENTS:
- 6+ years of software engineering experience with a track record of owning significant technical scope (e.g., driving projects from design through production or acting as de facto tech lead)
- Deep understanding of Kubernetes internals: controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns
- Solid grasp of distributed systems fundamentals—fault tolerance, graceful degradation, failure handling in large-scale environments
- Experience operating control planes and low-level components of large-scale Kubernetes clusters
- Experience with observability at scale: Prometheus, Grafana, distributed tracing, actionable alerting systems
- Strong programming skills in Go and Python; ability to collaborate on shared codebases
- Solid knowledge of Linux systems, networking, containers, and cloud infrastructure
- Pride in owning and delivering core product and platform components
PREFERRED QUALIFICATIONS:
- Experience building and operating managed Kubernetes services (GKE, EKS, AKS) or working on Kubernetes control plane components
- Hands-on experience with NVIDIA's GPU/networking ecosystem: GPU Operator, device plugins, DCGM, MIG, Network Operator, NCCL tuning
- Familiarity with HPC and traditional job schedulers (Slurm) and Kubernetes-native batch scheduling (KAI, Volcano, Kueue)
- Familiarity with GPU, InfiniBand, RDMA, or high-performance computing on Kubernetes
- Exposure to storage architecture for AI/ML workloads
- Past contributions to CNCF projects or Kubernetes SIGs