SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer - Core Cloud Platform

Lambda - San Francisco, CA, USA - Hybrid - posted 2026-07-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is building the superintelligence cloud, providing AI infrastructure to researchers, enterprises, and hyperscalers. As a Senior Software Engineer on the Core Cloud Platform team, you will design and operate the control-plane systems that power Lambda's GPU cloud infrastructure. You'll own core platform capabilities spanning compute lifecycle management, bare metal orchestration, maintenance workflows, deployment readiness, reliability, and operational tooling. Your work will directly translate physical GPU infrastructure into reliable, customer-facing cloud capacity through APIs, workflows, orchestration layers, schedulers, and operational surfaces. Key responsibilities include building and operating cloud platform services for compute lifecycle, bare metal hosts, capacity, and placement workflows; designing reliable APIs, backend services, state machines, and orchestration systems; implementing bare metal lifecycle systems (launch, terminate, restart, host reclaim, validation, quarantine, return-to-pool); improving deployment, observability, testing, alerting, and operational readiness for business-critical services; debugging complex production issues across distributed services and infrastructure; partnering with infrastructure, networking, fleet, security, and product teams; and contributing to architecture, design reviews, and team mentoring. You should have 6+ years of professional software engineering experience building production backend or distributed systems, strong proficiency in Python, Go, or similar languages, and experience designing and operating APIs, workflow engines, schedulers, or orchestration services. Deep understanding of reliability fundamentals (fault tolerance, idempotency, retries, state machines, failure handling) and production debugging is essential. Experience with cloud infrastructure primitives, Linux, containers, Kubernetes, and infrastructure automation is required. You should have owned production services, participated in on-call rotations, and improved systems based on operational learnings. Nice-to-have skills include experience with cloud control planes, compute platforms, bare metal, GPU infrastructure, HPC, Kubernetes, Slurm, host lifecycle management, provisioning, validation, firmware, BMC/Redfish, fleet management, Temporal, Airflow, event-driven systems, networking (VPCs, firewalls, SDN, InfiniBand), and security-minded engineering. You'll thrive in this role if you enjoy building foundational cloud systems where correctness and reliability matter, can work across product and infrastructure boundaries, break down ambiguous problems into concrete solutions, care deeply about production behavior, move fast while maintaining high standards, and communicate clearly about risks and tradeoffs. The position requires 4 days per week in the San Francisco/Fremont office, with Tuesday as the designated work-from-home day.

Similar roles