SlipstreamJobsFresh Startup & VC-Backed Jobs

Research Scientist, Scaling RL

Periodic Labs - Menlo Park, CA, United States - In-office - posted 2026-10-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 250,000 - 350,000 / annual

Periodic Labs is an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. The company is backed by world-class investors and operates at the pace the frontier requires. As a Research Scientist focused on Scaling RL, you will study how reinforcement learning scales with training compute and develop better algorithms, taking ideas from controlled experiments to the company's largest model runs like Periodic Neon. You'll work on frontier models to develop deep scientific knowledge and reasoning for scientific tasks. Key responsibilities include: - Design experiments to understand how RL performance scales with compute, model size, data, and reward quality, building on work such as ScaleRL - Develop better RL algorithms, spanning policy optimization, advantage estimation, exploration, and credit assignment for long-horizon RL tasks - Build adaptive sampling and curriculum methods that adjust task difficulty, problem selection, and the number of rollouts as models improve - Study bias and stability during RL training, including importance-sampling corrections and methods to tackle policy staleness and training–inference mismatch - Improve compute efficiency across training and inference through experiments with hyperparameters, such as length penalties, rollout counts, batch sizes, and update schedules You will work across a complex training stack to implement, debug, and test new research ideas, with a focus on designing small-scale RL setups that transfer to large-scale training runs. REQUIREMENTS: - 5+ years of experience - Bachelor's degree or equivalent experience - Hands-on experience training LLMs with reinforcement learning - Strong attention to detail and rigorous approach to answering questions scientifically - Ability to design small-scale RL setups that transfer to large-scale training runs - Comfort working across a complex training stack to implement, debug, and test new research ideas The company sponsors visas and will assist with the immigration process.

Similar roles