SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer - Storage Control Plane

Lambda - San Francisco, CA, USA - Hybrid - posted 2026-08-12

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is building the superintelligence cloud, providing AI infrastructure to researchers, enterprises, and hyperscalers. The Infrastructure Storage Team is seeking a Senior Software Engineer to design and build the next-generation storage control plane that powers Lambda's AI and machine learning infrastructure. In this role, you will architect a vendor-agnostic control plane that provisions, scales, heals, and meters storage across multiple platforms including VAST Data, WEKA, DDN, Pure, NetApp, Ceph, and MinIO. You'll define internal abstraction layers that hide vendor-specific APIs and failure semantics behind a single declarative interface, enabling new vendor integrations without re-architecture. Key responsibilities include building reconciliation-loop and CRD-based orchestration using Kubernetes controllers and operators to manage capacity, tenancy, encryption domains, and placement across data centers and availability zones. You'll own multi-tenant isolation end-to-end, including namespace partitioning, per-tenant QoS and rate limiting, credential lifecycle management, and blast-radius containment. You'll design the capacity and placement engine with awareness of PCIe topology, NUMA, and failure domains—critical for optimizing storage performance relative to GPU placement. You'll instrument the entire system with SLI/SLO definitions, fleet-wide performance regression detection, and observability pipelines that make petabyte-scale fleets debuggable. This is a hands-on engineering role requiring deep expertise in distributed systems, storage protocols (S3, NFS), and file systems like Ceph or DAOS. Required qualifications: Bachelor's or Master's in Computer Science; 5+ years developing storage systems; proven distributed systems programming experience; strong C, C++, Go, or Python skills; Linux kernel and system-level programming knowledge; familiarity with storage protocols and containerization (Docker, Kubernetes); CI/CD and QA practices for distributed systems. Nice-to-have skills include AI/ML workload experience, data center networking knowledge (InfiniBand, RoCE), storage performance tuning, GPU/DPU acceleration, production experience with enterprise storage platforms, and publications at top conferences (SNIA SDC, FAST, USENIX ATC).

Similar roles