SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Babylist is rebuilding its engineering culture around AI-first principles, where engineers own problems end-to-end with short feedback loops and rapid iteration. The Platform team is the foundation supporting 9 million+ users and all engineering teams across the organization.
As Staff SRE, you will own infrastructure and reliability practices that keep Babylist's systems fast, scalable, and dependable. You'll actively evolve AWS infrastructure, CI/CD systems, and developer tooling—this is not a maintenance role. Your decisions have wide leverage across all of Babylist Engineering as the company grows beyond its e-commerce and registry roots into health, media, mobile, and new product surfaces.
Key responsibilities include:
- Infrastructure ownership: manage and evolve the AWS environment using Terraform, keeping EKS clusters, databases, and core services current and performant
- CI/CD reliability: own the speed and reliability of CI systems for the entire Engineering org
- Developer support: unblock engineers fast when environments break across local, staging, and production
- Monitoring & alerting: establish best practices so the right people get paged for the right reasons
- Incident response: lead or support incident response, drive post-incident reviews, and prevent recurrence
- Platform strategy: contribute to architectural decisions shaping infrastructure evolution
You bring deep hands-on Terraform expertise, proven AWS experience at scale (EKS, RDS, networking, DNS, CDNs, load balancers), and production Kubernetes experience. You're comfortable designing and improving CI/CD systems (CircleCI, GitHub Actions, or similar), have strong observability instincts (Datadog, Sentry, PagerDuty, Cronitor), and are experienced with on-call and incident management. You naturally reach for AI in your work and stay curious about emerging tools. You support developers across all environments as a resource, not a gatekeeper.
The infrastructure is solid but actively evolving—you're shaping what comes next, not inheriting chaos. This staff-level role offers real cross-team visibility and influence over how Babylist engineers build and ship.