SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 250,000 - 350,000 / annual
Periodic Labs is an AI and physical sciences company building frontier models to accelerate breakthroughs in materials, energy, and scientific discovery. As a Midtraining Research Engineer, you will work on improving the scientific reasoning capabilities of large language models through data curation, synthetic data generation, and large-scale training experiments.
Your responsibilities will include identifying and processing novel sources of scientific data for model training, generating high-quality synthetic data to fill knowledge gaps, and building evaluations that correlate with downstream scientific task performance. You'll collaborate closely with RL researchers, physicists, and chemists to develop and apply advanced techniques such as self-distillation and on-policy distillation to enhance model capabilities.
You will design and execute large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs. A key part of your role involves building tools to investigate how data choices shape model intelligence and laying groundwork for pre-training efforts.
Ideal candidates have hands-on experience training LLMs on curated token mixes at scale, with prior mid-training or pre-training experience at big labs. You should be comfortable with self-distillation and on-policy distillation methods in production pipelines, able to calculate scaling laws and compute-optimal hyperparameters, and skilled across data, evals, and training infrastructure. Experience optimizing throughput and reliability for distributed training, background in AI for science or specialized domain data (protein, materials), and live-run eval tracking are strong additional qualifications.