SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer - Managed Kubernetes

Lambda - San Francisco, CA, USA - Hybrid - posted 2026-09-24

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is building the AI Cloud infrastructure of the future. As a Senior Software Engineer on the Managed Kubernetes (Mk8s) team, you will design and build the orchestration platform that powers mission-critical AI workloads at scale—think GKE, but purpose-built for AI and running on bare metal. You will work at the intersection of distributed systems, GPU-accelerated computing, and cloud-native infrastructure. Your responsibilities include: - Design, build, and maintain scalable control plane services, operators, and custom Kubernetes controllers; develop automation in Go/Python for end-to-end cluster lifecycle management (provisioning, upgrades, patching, deletion) - Build GPU-aware orchestration systems supporting GPU scheduling and resource allocation within the platform architecture - Partner with the Network team on networking solutions for AI workloads: CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and GPUDirect - Write resilient systems that handle failure gracefully—timeouts, retries, backoff, degraded-mode operation—across large-scale distributed environments - Develop platform services for inference: model serving infrastructure, autoscaling based on inference load, multi-model deployment patterns - Build internal tools and CLIs enabling ML/AI teams to deploy and monitor their own inference services - Support and debug production issues through on-call rotation Lambda serves tens of thousands of customers ranging from AI researchers to enterprises and hyperscalers. The company was founded in 2012, has 500+ employees, and is backed by notable investors including NVIDIA, Andrej Karpathy, ARK Invest, and In-Q-Tel. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. This role requires presence in the San Francisco, San Jose, or Bellevue office 4 days per week; Tuesday is the designated work-from-home day. REQUIREMENTS: - 6+ years of software engineering experience with a track record of owning significant technical scope (e.g., driving projects from design through production or acting as de facto tech lead) - Deep understanding of Kubernetes internals: controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns - Solid grasp of distributed systems fundamentals—fault tolerance, graceful degradation, failure handling in large-scale environments - Experience operating control planes and low-level components of large-scale Kubernetes clusters - Experience with observability at scale: Prometheus, Grafana, distributed tracing, actionable alerting systems - Strong programming skills in Go and Python; ability to collaborate on shared codebases - Solid knowledge of Linux systems, networking, containers, and cloud infrastructure - Pride in owning and delivering core product and platform components PREFERRED QUALIFICATIONS: - Experience building and operating managed Kubernetes services (GKE, EKS, AKS) or working on Kubernetes control plane components - Hands-on experience with NVIDIA's GPU/networking ecosystem: GPU Operator, device plugins, DCGM, MIG, Network Operator, NCCL tuning - Familiarity with HPC and traditional job schedulers (Slurm) and Kubernetes-native batch scheduling (KAI, Volcano, Kueue) - Familiarity with GPU, InfiniBand, RDMA, or high-performance computing on Kubernetes - Exposure to storage architecture for AI/ML workloads - Past contributions to CNCF projects or Kubernetes SIGs

Similar roles