SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 144,000 - 240,000 / annual
Lila Sciences is building Scientific Superintelligence to accelerate scientific discovery with AI. As a Senior Data Engineer, you'll design and build ETL pipelines and data models for Lila's scientific data platform, working at the intersection of data engineering, computational biology, chemistry, and materials science.
Your core responsibility is transforming raw lab instrument outputs into validated, analysis-ready datasets. You'll tackle complex data modeling challenges: converting messy, per-instrument measurements into clean, well-typed data that is efficient to query, reliable to use, and ready for downstream AI and scientific analysis. You'll partner closely with AI researchers and experimentalists to understand domain requirements and deliver solutions that accelerate their work.
Key responsibilities include: designing pipelines that turn raw lab output into analysis-ready scientific data; modeling heterogeneous data from biological, chemistry, and materials instruments; building validation checks, schema-evolution gates, and data quality workflows; developing reusable analysis functions for scientific and AI research workflows; improving automation and observability across instrument-to-result data flows; and building canonical datasets that scientists and AI researchers can trust.
You should have 2–6 years of experience in data engineering, bioinformatics, cheminformatics, or computational science. Required skills include strong Python (typed, tested, production-quality code), strong SQL (Postgres or similar), ETL pipeline and data model design, data science fundamentals (statistics, pandas, NumPy), experience translating noisy scientific measurements into validated datasets, and workflow orchestration experience (Flyte, Airflow, Prefect, Dagster, or Nextflow). Active use of AI coding tools in day-to-day work is expected.
Bonus experience includes columnar/lakehouse stacks (Parquet, Iceberg, DuckDB, Polars, Ibis), event-driven pipelines (NATS, Kafka), lab instrument data formats/LIMS/ELN systems, life sciences assays (sequencing, imaging, flow cytometry), materials/chemistry methods (XRD, XRF, SEM, TGA, DSC), and curve fitting/peak detection techniques.