SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Chai Discovery builds a design suite for molecules using frontier AI models that learn biochemical structure and interaction. The company partners with leading pharmaceutical companies including Eli Lilly, Pfizer, and Novartis to power drug discovery programs.
As a Research Engineer on the ML Infrastructure team, you will own the distributed systems that enable the company's models to train and run performantly, reliably, and at scale. Your responsibilities include:
- Architecting, debugging, and optimizing the distributed ML training stack across model, layer, and kernel levels, eliminating runtime and reliability bottlenecks on large GPU clusters.
- Profiling end-to-end training runs to identify bottlenecks across compute, communication, and storage, and building tooling to monitor throughput, utilization, and uptime.
- Optimizing ML workloads through parallelism strategies, quantization, and custom CUDA/Triton kernels.
- Collaborating closely with Research Scientists to ensure new model architectures and training recipes scale efficiently from early experiments to frontier-scale runs.
- Owning reliability of the training stack, including fault tolerance, checkpointing, and deterministic orchestration for long-running, large-scale jobs.
You will work at the intersection of AI research and frontier biology with a mission-driven team focused on high velocity and ownership.
REQUIREMENTS:
- 4+ years of industry experience working within AI/ML infrastructure teams
- Proficiency in Python and PyTorch or JAX
- Strong software systems design skills, with comfort operating across the stack from model code down to kernels
- Experience with orchestrating GPU clusters and large-scale model training
- Experience with optimizing ML workloads: parallelism, quantization, CUDA/Triton kernels