SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lambda is building AI cloud infrastructure and seeks a Staff Software Engineer to lead the development of a managed Kubernetes platform purpose-built for AI workloads running on bare metal. This is a foundational technical leadership role shaping the infrastructure powering next-generation AI training and inference at scale.
In this role, you will drive the technical vision for Lambda's managed orchestration services, including Managed Kubernetes, Managed Slurm on Kubernetes, and higher-level platform services for inference and AIOps. You'll work at the intersection of distributed systems, GPU-accelerated computing, and cloud-native infrastructure to build reliable, performant, and elegant systems for customers ranging from AI researchers to enterprises and hyperscalers.
Key responsibilities include:
**Product Engineering:** Drive technical vision for the bare-metal Kubernetes platform covering control plane scalability, multi-tenancy, cluster lifecycle management, and high availability. Integrate and extend NVIDIA's open-source ecosystem (GPU Operator, Network Operator, DCGM, NCCL, AICR, Topograph). Design GPU-aware orchestration systems and lead development of managed services. Inform networking solutions for AI workloads including CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and GPUDirect. Partner with Storage teams on architecture requirements. Build the foundation for Managed Slurm on Kubernetes. Design higher-level platform services for inference including model serving, autoscaling, and multi-model deployment. Design self-healing systems and automation for incident response and platform resilience. Lead chaos engineering efforts and establish operational excellence for managed services.
**Cross-Functional Infrastructure Leadership:** Serve as technical bridge between Orchestration and other infrastructure teams (Network, Storage, Security). Drive infrastructure-wide decisions enabling successful managed services. Provide input on bare-metal provisioning, network topology, and storage systems. Champion consistency and standardization across Lambda's infrastructure stack. Work directly with customers and internal teams to understand deployments and chart paths to the managed platform.
**Technical Leadership:** Set technical direction for Kubernetes services, influencing roadmap and prioritization. Drive design reviews and sessions ensuring scalable, maintainable systems aligned with customer needs. Mentor and grow engineers, establishing best practices. Collaborate cross-functionally. Engage with NVIDIA and open-source community. Represent Lambda externally through technical content and customer engagements. Shape AIOps vision for automated capacity planning, anomaly detection, and predictive maintenance.
Lambda was founded in 2012 and has 500+ employees. The company is backed by notable investors including NVIDIA, Andrej Karpathy, ARK Invest, In-Q-Tel, and others. Lambda has research papers accepted at top ML and graphics conferences including NeurIPS, ICCV, SIGGRAPH, and TOG.
**Requirements:**
- 10+ years of software engineering, platform engineering, or SRE experience, with at least 5 years focused on Kubernetes at scale
- Expert-level understanding of Kubernetes internals: API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns
- Holistic infrastructure expertise across compute, networking, storage, and security—not just Kubernetes in isolation
- Strong software engineering skills in Go (required) and Python; production-quality code
- Deep experience with GPU orchestration in Kubernetes: NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, GPU-aware scheduling. Familiarity with NVIDIA Network Operator and GPUDirect strongly preferred
- Proven track record of technical leadership: driving design decisions across teams, mentoring engineers, influencing infrastructure direction
- Deep experience designing and operating managed services or multi-tenant platforms
- Strong understanding of distributed systems principles: consensus, fault tolerance, consistency models, graceful degradation
- Experience with observability at scale: Prometheus, Grafana, distributed tracing, actionable alerting
- Solid knowledge of Linux systems and networking (L2-L7), including high-performance networking concepts (RDMA, InfiniBand, RoCE)
- Experience with infrastructure-as-code and GitOps workflows
**Preferred Qualifications:**
- Experience building and operating managed Kubernetes services (GKE, EKS, AKS) or working on Kubernetes control plane components
- Hands-on experience with NVIDIA's open-source ecosystem beyond GPU Operator: Network Operator, NCCL tuning, Topograph, AICR
- Familiarity with HPC and traditional job schedulers (Slurm) and Kubernetes-native batch scheduling (KAI, Volcano, Kueue)
- Background in confidential computing
- Experience migrating customers or workloads from legacy infrastructure to standardized platforms
- Contributions to CNCF projects, Kubernetes SIGs, or NVIDIA open-source projects
- Familiarity with security and compliance in multi-tenant environments: RBAC, Pod Security Standards, network policies, workload isolation
- Background in ML infrastructure: training clusters, inference serving, simulation