SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer - Core Cloud Platform

Lambda - San Francisco, CA, USA - Hybrid - posted 2026-07-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is an AI cloud infrastructure company serving tens of thousands of customers from researchers to enterprises and hyperscalers. The Core Cloud Platform team powers compute provisioning and infrastructure orchestration across Lambda's physical data centers. As a Senior Site Reliability Engineer, you will improve the reliability, scalability, and operational maturity of these systems as Lambda's fleet and customer base grow. You will own the operation and scaling of critical platform services across data centers, working with Kubernetes, infrastructure automation, observability, deployment systems, and incident response. Key responsibilities include: improving reliability of compute provisioning and instance lifecycle management; building comprehensive monitoring, alerting, and tracing for service health and customer-impacting failures; defining SLIs, SLOs, and error budgets; automating detection and remediation of configuration drift and failed workflows; designing fault-isolation mechanisms to reduce blast radius; leading production incident response and postmortems; and mentoring engineers to raise reliability standards across the organization. You will partner closely with Compute, Networking, Storage, Security, and Support teams, participate in on-call rotations, and drive automation to improve on-call sustainability. The role requires 7+ years in site reliability, infrastructure, distributed systems, or production engineering. Essential qualifications include deep production Kubernetes experience, understanding of Kubernetes architecture and failure modes, experience with physical data centers or private cloud environments, proficiency with Terraform or similar IaC tools, CI/CD/GitOps workflow experience (Argo CD, Flux, Helm, Kustomize), observability platform expertise (OpenTelemetry, Prometheus, Grafana, Datadog), and ability to build production tooling in Go, Python, or similar languages. You should understand distributed systems concepts, have SLI/SLO operational experience, lead effectively during high-severity incidents, and approach operational issues as engineering problems. Nice-to-have skills include AI infrastructure or GPU platform experience, multi-region/data center operations, Kubernetes controllers and operators, etcd expertise, Linux systems knowledge, chaos engineering, and compliance framework familiarity.

Similar roles