SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer - Managed Kubernetes

Lambda - Bellevue, WA, United States - Hybrid - posted 2026-08-06

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is building the AI Cloud of the future and seeks a Staff Engineer to lead development of its Managed Kubernetes platform—a purpose-built, bare-metal orchestration system for AI workloads comparable to GKE but optimized for GPU-accelerated computing. In this foundational technical leadership role, you will shape the infrastructure powering next-generation AI training and inference at scale. As a Staff Engineer on the Orchestration team, you'll drive technical vision for managed orchestration services including Managed Kubernetes, Managed Slurm on Kubernetes, and higher-level platform services for inference and AIOps. You'll work at the intersection of distributed systems, GPU-accelerated computing, and cloud-native infrastructure to build reliable, performant, and elegant systems. Key responsibilities include: **Product Engineering:** Drive technical vision for the Managed Kubernetes bare-metal platform, including control plane scalability, multi-tenancy, cluster lifecycle management, and high availability. Integrate and extend NVIDIA's open-source ecosystem (GPU Operator, Network Operator, DCGM, NCCL, AICR, Topograph). Design GPU-aware orchestration systems and lead development of managed services. Collaborate with Network and Storage teams on CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, GPUDirect, and storage architecture for AI workloads. Build foundations for Managed Slurm on Kubernetes. Design higher-level platform services for inference, including model serving, autoscaling, and multi-model deployment. Lead chaos engineering efforts and establish operational excellence for managed services (upgrade automation, security patching, zero-downtime maintenance). **Cross-Functional Infrastructure Leadership:** Serve as technical bridge between Orchestration and other infrastructure teams (Network, Storage, Security). Drive infrastructure-wide decisions enabling successful managed services. Provide input on bare-metal provisioning, network topology, and storage systems. Champion consistency and standardization across Lambda's infrastructure stack. Work directly with customers and internal teams to understand deployments and chart paths to the managed platform. **Technical Leadership:** Set technical direction for Kubernetes services, influencing roadmap and prioritization. Drive design reviews and sessions ensuring scalability and maintainability. Mentor engineers and establish best practices for Kubernetes development and distributed systems. Collaborate cross-functionally. Engage with NVIDIA and the open-source community. Represent Lambda externally through technical content and customer engagements. Shape AIOps vision for automated capacity planning, anomaly detection, and predictive maintenance. You are a creative, innovative engineer operating at high velocity who finds elegant solutions and ships quickly. You embrace modern tools and AI-assisted development to accelerate productivity and multiply impact. You're energized by building new things. Required: 10+ years software engineering, platform engineering, or SRE experience, with at least 5 years focused on Kubernetes at scale. Expert-level understanding of Kubernetes internals (API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, extension patterns). Holistic infrastructure expertise synthesizing knowledge across compute, networking, storage, and security.

Similar roles