SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 185,500 - 232,000 / annual
Formation Bio is an AI-driven pharmaceutical company focused on accelerating drug development and clinical trials through technology platforms and AI. The company was founded in 2016 (originally as TrialSpark Inc.) and is backed by leading investors including a16z, Sequoia, Sanofi, Thrive Capital, and others.
As a Senior Site Reliability Engineer, you will own the infrastructure and operational platform that enables Formation Bio's engineering organization to ship reliable software quickly and safely. You will work across cloud infrastructure, developer platforms, observability, and production workloads—including product applications, internal tools, data systems, and ML/AI workloads. Formation Bio is an AI-native engineering organization, and you will use modern AI tools, including agentic coding systems, as part of your daily practice while applying strong judgment to validate their output and maintain reliable production systems.
Key responsibilities include:
- Owning infrastructure and operational platform for shared engineering workloads (compute, runtime environments, orchestration, deployment, observability, access controls, reliability)
- Building and operating secure, observable, reliable infrastructure for product applications, containerized services, internal tools, data systems, ML pipelines, inference, and agentic software
- Researching, developing, and maintaining core AWS infrastructure and cloud outposts for development, staging, and production environments
- Creating, reviewing, maintaining, and optimizing infrastructure as code, CI/CD pipelines, and reusable platform patterns
- Establishing strong operational practices including SLOs, monitoring, alerting, runbooks, incident response, advanced diagnostics, and root cause analysis
- Partnering with Product Engineering, Data Engineering, and Data Science teams to evaluate and implement architecture and infrastructure for product software, data systems, model training, and inference
- Using AI tools to accelerate infrastructure development, investigate incidents, improve documentation, and build automation
- Participating in support rotation and incident response with a strong bias toward automation
- Writing and reviewing requirements, design documents, and operating procedures; mentoring engineers on infrastructure and SRE fundamentals
The role is hybrid, requiring 3 days per week in office. Primary hiring focus is New York City and Boston metro areas, with consideration for Research Triangle (NC) and San Francisco Bay Area candidates.
Requirements:
- 5+ years of relevant experience in Site Reliability Engineering, infrastructure, systems, DevOps, or similar discipline
- Production experience operating cloud infrastructure and distributed systems with strong operational and reliability judgment
- Experience with advanced diagnostics, incident response, root cause analysis, observability, and automation
- Experience with AWS and Snowflake (required); Azure, GCP, and/or Vercel experience is a plus
- Working experience with Docker, GitHub, Kubernetes, Python, Terraform or OpenTofu, and virtual networking (Terragrunt is a plus)
- Experience supporting production ML or AI workloads, MLOps infrastructure, workflow orchestration, model serving, or related platform is a plus; strong SRE and infrastructure fundamentals are core requirements
- Experience managing shared COTS and FOSS software applications in addition to internal software and tools
- Daily fluency with AI tools, including LLMs and agentic coding systems, paired with strong engineering judgment and high bar for validation
- Exceptional collaboration and communication skills across engineering, Data Science, Data Engineering, Security, and non-technical partners
- Experience operating infrastructure in a regulated or validated environment is a plus