SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior ML Engineer – ADMET & Toxicity Networks

Apheris - Remote - Remote - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Apheris is building AI infrastructure for pharmaceutical R&D, hosting federated data networks for drug discovery. The company enables leading pharma teams to discover and develop drugs faster by training models on proprietary industry datasets while keeping data control and IP protected. You will own the training and evaluation pipelines behind Apheris's ADMET and toxicity networks—the machinery that turns partner data into trained, benchmarked, released models. This is a hands-on engineering role at the intersection of molecular ML and federated learning. Your code runs inside partner environments on data you cannot see, alongside scientists at major pharma companies. You'll work closely with the scientific lead for toxicity, who owns what the models predict and why; you own how it gets built, validated, evaluated, and shipped reliably enough that partners will stake drug programs on the results. The ADMET network is live and expanding into toxicity, and you'll build the pipelines that expansion runs on. Key responsibilities: - Own model pipelines end-to-end: take partner data from landing through to released, benchmarked model weights, including data preparation, training, evaluation, and release for ADMET and toxicity endpoints. - Build validation that partners run alongside Apheris: schema and data contracts, validators that return actionable errors, and QC/profiling reports that partners can act on without exposing raw data. - Make federated runs reproducible and auditable: versioned configs, pinned data snapshots, and provenance for every released model so the exact source of any weights can be documented. - Build evaluation that survives scrutiny: leakage-safe splitting, held-out benchmarks, and honest performance reporting so numbers presented to partners hold up under review. - Harden research into product: turn prototypes and research code into tested, modular systems and hand them cleanly to engineering for scaling into Foundry. - Work across the boundary: translate scientific requirements into pipeline behavior with the science team and surface data or modeling risks early to partners and internally. Requirements: - 5+ years building ML systems in Python with genuine software engineering discipline: version control, tested modular code, code review, and interfaces others can use. - Hands-on molecular ML or cheminformatics experience: RDKit, fingerprints and descriptors, or graph/transformer models applied to property, activity, or toxicity prediction. - Experience building training and evaluation pipelines that other people run, not one-off notebooks. - Understanding of how molecular ML fails: data leakage, split design, applicability domain, dataset shift, and over-optimistic benchmarks—and ability to design against them by default. - Ability to write validators and data contracts that hold up under partial visibility, where you cannot inspect the data yourself. - Comfort with PyTorch or equivalent modern ML stack. - Strong ability to work with scientists: translate ambiguous scientific requirements into defined, testable pipeline behavior. Nice-to-have: - Federated learning, privacy-preserving ML, or other multi-party training environments. - ML Ops or ML infrastructure experience, particularly Kubernetes-based training, evaluation, or deployment workflows. - Production-grade model delivery in regulated, enterprise, pharmaceutical, or biotech settings. - Familiarity with public ADMET, toxicity, and bioactivity data resources (ChEMBL, Tox21, ToxCast) and their gotchas. - Open-source contributions or publication record in cheminformatics, molecular ML, or applied machine learning.

Similar roles