SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Babylist is rebuilding its engineering culture around AI-first principles, and the Platform team is the foundation that every engineering team depends on. As a Staff SRE, you'll own infrastructure and reliability practices supporting 9 million+ users across e-commerce, registry, health, media, and mobile products.
You'll be responsible for evolving AWS infrastructure using Terraform, managing EKS clusters, RDS databases, and core services. You'll own CI/CD reliability across the entire engineering organization—every deploy starts with your systems. You'll establish monitoring and alerting standards, lead incident response, and drive post-incident reviews to prevent recurrence. You'll be the resource engineers turn to when environments break, unblocking them quickly across local, staging, and production.
The role requires deep hands-on expertise: Terraform IaC ownership, proven AWS experience at scale (EKS, RDS, networking, DNS, CDNs, load balancers), production Kubernetes operations, CI/CD system design (CircleCI, GitHub Actions), and strong observability practices (Datadog, Sentry, PagerDuty). You've run on-call rotations, debugged hard infrastructure problems, and actually changed things after post-mortems.
This is a staff-level position with real cross-team visibility and leverage. You'll influence how Babylist engineers build and ship, contribute to multi-year architectural decisions, and work directly with product, design, and business partners. The infrastructure is solid but actively evolving—you're shaping what comes next, not inheriting chaos. AI tools are natural to the workflow; you should already be using AI daily to move faster and improve output.