SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer - Managed Kubernetes

Lambda - San Francisco, CA, USA - Hybrid - posted 2026-08-14

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is a leader in AI cloud infrastructure serving tens of thousands of customers globally, from AI researchers to enterprises and hyperscalers. The company's mission is to make compute as ubiquitous as electricity and democratize access to superintelligence. As a Senior Site Reliability Engineer, you will operate and maintain bare-metal Kubernetes clusters at scale, managing thousands of nodes across Lambda's cloud platform. Your responsibilities include cluster degradation and recovery, incident response using fleet management tools, and participation in a well-managed on-call rotation for critical incidents. You'll design and build scalable control plane services, operators, and custom controllers for Kubernetes, while developing automation for the complete cluster lifecycle—provisioning, upgrades, patching, and deletion. You'll use Python and Golang to create tooling, automate platform quality validation, and define SLOs/SLIs for Kubernetes services and workloads. Collaboration is central to the role: you'll work closely with HPC Ops and Datacenter Ops teams on low-level and cross-functional issues, assist customers with Kubernetes integration and authentication questions, and participate in incident response via tickets, messaging, or calls. You bring 6+ years of SRE or operations engineering experience with deep Linux cluster and systems knowledge. You're proficient in Go and Python, with hands-on experience operating Kubernetes in production (on-prem, EKS, GKE, or similar). You're comfortable with GitOps tools like ArgoCD, Helm, Kubernetes operators, and observability platforms such as Prometheus and Grafana. Experience provisioning Kubernetes via kubeadm, Cluster API, or similar tools is essential. Nice-to-have qualifications include deep Kubernetes expertise (CRDs, CSI, CNI, operator coding), exposure to HPC clusters and AI/ML workloads, hybrid/multi-cloud Kubernetes experience, and contributions to CNCF projects or Kubernetes SIGs. Lambda offers generous cash and equity compensation, comprehensive health/dental/vision coverage, wellness and commuter stipends, 401k matching, and flexible PTO. The role requires presence in the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday as the designated work-from-home day.

Similar roles