SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 268,000 - 358,000 / annual
Lila Sciences is building Scientific Superintelligence to solve major challenges by combining automated large-scale data generation with AI. The Life Science AI team develops machine learning systems for automated reasoning on biological data, merging state-of-the-art ML with breakthrough biology.
You will focus on domain models for perturbation biology, genetics, and high-dimensional experimental readouts that connect biological mechanism, experimental intervention, and therapeutic opportunity. A key differentiator at Lila is that data is generated for the model—the company runs its own experiments at scale, and you will help decide what gets measured. The role bridges the gap between what to learn from fixed datasets and what datasets should exist.
Key responsibilities include:
- Build domain models for perturbation, genetic, and multimodal experimental data; translate biological questions into rigorous ML problem formulations with interpretable structure where biology supports it
- Make uncertainty a deliverable through calibrated posteriors, honest error bars, and outputs that communicate prediction reliability, especially when predicting into unmeasured conditions
- Pre-register and beat baselines; simple entity-mean, additive, and linear baselines are stated before any model is fit
- Partner with experimental scientists to guide data generation and model validation, shaping what gets measured and at what precision; build experiment-selection methods to reduce uncertainty where it matters
- Design benchmarks connecting model performance to biological and therapeutic consequence; apply model outputs to prioritize targets and mechanisms
- Support integration of domain models into agentic workflows; deploy tools and collaborate on creating environments to train reasoning models that drive agents
- Communicate findings clearly to technical and cross-functional audiences including scientists, engineers, product partners, and therapeutic stakeholders
- Support external scientific visibility through publications, presentations, and engagement with the ML/AI for Biology and computational biology communities
Model outputs are reasoned over by systems and scientists choosing what to run next and making mechanistic calls downstream. Both need to know why, not only what. A model that returns a quantity a biologist can argue with, models that transfer to new contexts, compose into downstream calculations, and can be checked against independent measurements are often worth more than a more accurate one that returns an embedding.
REQUIREMENTS:
- PhD in machine learning, statistics, computational biology, computer science, bioengineering, physics, or a related quantitative field, with a strong publication record or equivalent industry impact
- Experience developing models for high-dimensional biological data, with judgment to connect ML methods to biological mechanism and experimental design
- Generalization under structured sparsity: track record of building models that predict into conditions not directly observed (sparse or unbalanced experimental designs, held-out combinations, transfer to new contexts) rather than interpolating within densely sampled corpora
- Uncertainty and evaluation rigor: comfort with calibration, proper scoring rules, and evaluation design in work where decisions were made based on your numbers
- Strong programming skills, reliable ML research workflows, and clear communication across ML, biology, and experimental teams
- Track record of leading ambiguous research problems from formulation through execution
BONUS QUALIFICATIONS:
- Mechanistic and probabilistic modeling (strongest differentiator): Bayesian hierarchical models, simulation-based or likelihood-free inference, amortized posterior inference, state-space or ODE-based models, neural differential equations, or mechanism-informed ML
- Experience with pharmacokinetic/pharmacodynamic, systems-biology, or physical modeling where parameters had to mean something
- Familiarity with active learning, Bayesian optimal experimental design, closed-loop experimentation, or lab-in-the-loop systems
- Experience deploying research models into scientific decision-making workflows, including serving models as tools other systems call