SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 185,500 - 232,000 / annual
Formation Bio is a tech and AI-driven pharma company founded in 2016 (originally TrialSpark Inc.) that accelerates drug development and clinical trials through advanced technology platforms. The company partners with pharma companies, research organizations, and biotechs to develop drug candidates past clinical proof of concept, backed by investors including a16z, Sequoia, Sanofi, Thrive Capital, and Spark Capital.
As a Senior Data Engineer, you will build trusted data systems supporting clinical operations, drug asset evaluation, business development, analytics, and machine learning through AI-enabled employees and agents. You will work across clinical, operational, and third-party data sources to design and operate reliable ingestion pipelines, transformations, data models, and data products. This role sits at the intersection of Product Engineering, Data Engineering, and Data Science.
Key responsibilities include:
- Design and operate production data systems ingesting clinical, operational, and vendor data into reliable, queryable data products
- Own shared canonical data models, data contracts, transformations, orchestration, warehouse models, and downstream interfaces
- Partner with Product Engineering on application data models, source-system contracts, APIs, events, and data access patterns
- Partner with Data Science on productionized training datasets, feature pipelines, data interfaces, and ML use cases
- Turn recurring data cleaning, normalization, and transformation work into versioned, tested, observable, maintainable production pipelines
- Build data products for clinical operations, asset evaluation, business development, analytics, machine learning, and AI-enabled systems
- Design data products that are semantically clear, discoverable, machine-readable, permission-aware, traceable, and safe to query
- Establish strong data quality, testing, freshness, completeness, lineage, documentation, and observability practices
- Own data governance practices for sensitive and regulated data, including access controls, auditability, and traceability
- Participate in support and incident response for data platform issues
- Use AI tools including LLMs and agentic coding systems to accelerate development while validating output
- Contribute to architecture reviews, mentor engineers, and improve engineering practices
Requirements:
- 5+ years of relevant data engineering experience building and operating production data systems
- Experience with pharmaceutical, biology, HIPAA, or other regulated data core to biotech (required)
- Strong Python and SQL skills with deep experience in data modeling and warehouse systems, especially Snowflake
- Experience with Dagster as orchestration system (or equivalent) and transformation tooling such as dbt
- Experience with data contracts, schema evolution, data quality testing, observability, lineage, and production incident response
- Experience integrating messy clinical, operational, vendor, or complex source data
- Working knowledge of Docker, GitHub, and Terraform or OpenTofu sufficient to partner with SRE
- Experience building data products and access patterns for applications, Data Science, analytics, human users, and AI-enabled systems
- Strong judgment about when to build reusable platform capabilities versus one-off solutions
- Daily fluency with AI tools and ability to validate generated code, transformations, and data-modeling decisions
- Exceptional collaboration and communication skills across Product Engineering, Data Science, Clinical Operations, Data Management, Business Development, and non-technical partners
- Experience working within and building validated computerized systems (CSV) is a plus