SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lambda is seeking a Staff Software Engineer to lead the technical vision for its next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges distributed systems and semiconductor architecture to enable reliable cloud provisioning and lifecycle management of heterogeneous compute platforms at massive scale.
You will provide hands-on technical leadership guiding development of a resilient compute control plane utilizing durable execution concepts and deep hardware integration. The position requires deep understanding of the entire stack: BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, and large-scale cloud service provider operations.
Key responsibilities include:
- Designing and implementing highly available and reliable GPU and CPU host and instance lifecycle control planes
- Guiding technical decisions on semiconductor architecture, BIOS/firmware settings, system boot methodologies, and DPU utilization
- Guiding design of compute platform multi-tenant security models
- Providing technical leadership and mentorship for senior engineers across multiple teams
- Collaborating with product and data center organizations to translate customer requirements into scalable infrastructure capabilities
- Working with customers to translate technical requirements into concrete engineering deliverables
- Setting engineering standards and leading design reviews for mission-critical cloud software at scale
Required qualifications: 10+ years experience with compute control plane distributed systems for deploying and lifecycle managing heterogeneous compute platforms in data centers; deep expertise in durable execution models and distributed systems for cloud provisioning; basic knowledge of software-defined networking; proven track record leading large-scale semiconductor hardware enablement initiatives; proven experience deploying net-new data centers into global compute platforms; proficiency in C/C++, Rust, Python, or Go.
Nice-to-have skills include knowledge of Nvidia AI Factory components, DOCA software, Linux kernel internals, virtualization technologies, Kubernetes, and high-performance networking protocols.
Lambda is an AI cloud infrastructure leader serving tens of thousands of customers from AI researchers to enterprises and hyperscalers. Founded in 2012 with 500+ employees, backed by notable investors including NVIDIA, ARK Invest, and In-Q-Tel. The role requires 4 days per week in office (Bellevue, San Francisco, or San Jose), with Tuesday as the designated work-from-home day.