SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer - Fleet Orchestration

Lambda - San Francisco, CA, USA - Hybrid - posted 2026-08-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is a leader in AI cloud infrastructure serving tens of thousands of customers globally, from AI researchers to enterprises and hyperscalers. The company's mission is to make compute as ubiquitous as electricity and democratize access to superintelligence. In this role, you will own the design and implementation of fleet data systems that make Lambda's GPU datacenter hardware legible at scale. Your responsibilities include: - Building and maintaining fleet data systems that index, validate, and reconcile physical host state against intended logical configuration across large-scale GPU environments - Leading the design and implementation of automation that orchestrates GPU cluster deployments from logical design import through racking, OS provisioning, validation, and customer hand-off - Designing systems that continuously validate consistency between intended and actual state, catching drift before it causes deployment failures or delays - Implementing ownership, global locking, readiness gating, and action safety systems that keep fleet operations coordinated across teams and tools - Owning technical contributions to new datacenter site bring-up, working cross-functionally with HPC Deployments, DC Ops, Network Engineering, and Core Infrastructure teams - Monitoring production health, maintaining SLAs, and driving resolution when systems drift You bring 5+ years of engineering experience with fluency in Python, Go, or similar languages. You're comfortable with APIs, distributed systems, and automation pipelines. You can reason about PXE boot, firmware provisioning, IPMI/BMC interfaces, and network services like DNS and DHCP from first principles. You have owned production systems with real SLAs and can lead technical design on medium-to-large features—taking ambiguous problems, writing design documents, driving alignment, and shipping solutions. You've worked cross-functionally to influence technical decisions beyond your immediate team and have mentored peers, leaving systems and teammates better than you found them. Nice-to-have qualifications include experience in ML/AI infrastructure, familiarity with datacenter physical infrastructure (racks, switches, InfiniBand fabric, power domains), a background blending software engineering with systems or infrastructure engineering, and experience with network source-of-truth systems like NetBox or DCIM tooling. Lambda was founded in 2012 and now has 500+ employees. The company is backed by notable investors including NVIDIA, Andrej Karpathy, ARK Invest, In-Q-Tel, and others. The role requires presence in the San Francisco or San Jose office 4 days per week, with Tuesday as the designated work-from-home day.

Similar roles