SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 250,000 - 350,000 / annual
Periodic Labs is an AI and physical sciences company building frontier models to accelerate breakthroughs in materials, energy, and beyond. The company is backed by world-class investors and operates at the pace required for cutting-edge research.
As a Research Scientist/Research Engineer focused on Midtraining, you will work on training frontier models to develop deep scientific knowledge and reasoning for scientific discovery. Your core responsibilities include:
- Identifying, processing, and curating novel sources of scientific data for large-scale model training
- Generating high-quality synthetic data to fill gaps in scientific knowledge and reasoning
- Building evaluations that correlate with downstream scientific task performance, collaborating with RL researchers, physicists, and chemists
- Developing and applying techniques such as self-distillation and on-policy distillation to improve model capability
- Designing and running large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs
- Building tools to investigate how data choices shape model intelligence
You will work at the intersection of data curation, evaluation design, and training infrastructure, directly contributing to the development of scientific reasoning capabilities in frontier models.
REQUIREMENTS:
- Experience training LLMs on curated mixes of trillions of tokens
- Experience on a dedicated evals team supporting a large production training run
- Hands-on use of self-distillation, on-policy distillation, or similar methods in a real training pipeline
- Experience with scaling laws and compute-optimal hyperparameters
- Comfort working across data, evals, and training infrastructure
- Minimum education: Bachelor's degree or equivalent experience
STRONG ADDITIONAL QUALIFICATIONS:
- Experience optimizing throughput and reliability for large-scale distributed training runs
- Background in AI for science or training on specialized domain data (e.g., protein, materials, or other scientific datasets)
- Experience creating evals or synthetic data for non-verifiable tasks and tracking performance over live runs
Visa sponsorship is available.