SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer - Storage

Lambda - San Francisco, CA, USA - Hybrid - posted 2026-08-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is a leader in AI cloud infrastructure serving tens of thousands of customers from AI researchers to enterprises and hyperscalers. The Storage Engineering team operates the backbone of Lambda's data platform, managing the full spectrum from low-level storage systems to customer-facing APIs and tooling. This role focuses on ensuring reliability and performance of storage systems operating at scale across Lambda's data centers. You will own the reliability, performance, and capacity health of Lambda's production storage fleet across all data centers, operating behind Lambda's software-defined data plane. Key responsibilities include building and maintaining monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures; investigating and resolving storage-related incidents using deep telemetry and performance profiling; automating ticketing, escalation, and incident-response workflows; and designing self-healing automation for common failure modes like drive replacement and node swaps. You'll implement CI/CD pipelines for storage automation, partner with Storage Engineers and Fleet Orchestration teams to automate deployment and configuration of software-defined storage across sites using tools like Ansible and Jenkins, and work with hardware and networking teams to diagnose low-level I/O and network issues. You'll participate in an on-call rotation with a focus on driving down mean time to recovery and building automation to prevent recurring incidents. Required qualifications include 5+ years operating Linux systems in production or HPC environments with hands-on storage experience at scale on scale-out or software-defined platforms (CEPH, Lustre, GPFS, or similar); hands-on experience operating Software-Defined Storage platforms at scale; strong incident-response instincts; working experience with monitoring platforms like Prometheus, Grafana, Datadog, or SumoLogic; Kubernetes and GitOps tooling experience; CI/CD tooling, containerization, and systems programming in Python or Go; Infrastructure as Code tools like Terraform and Ansible; and solid understanding of storage protocols across file, object, block, and structured storage. Nice-to-have skills include experience with cutting-edge SDS solutions like VAST or Weka, enterprise storage expertise, Kubernetes CSI driver experience, SR-IOV and virtualization knowledge, GPUDirect Storage and RDMA experience, NIC-level diagnostics, and open-source storage contributions.

Similar roles