SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Preference Model is building automated ML research engineering to advance frontier AI models. The company is developing high-quality RL training environments that reflect real-world complexity, with diverse tasks and robust reward functions. The founding team includes former members of Anthropic's data team who built infrastructure and datasets behind Claude.
In this role, you will work at the intersection of research and engineering to push the boundaries of post-training on large language models. You'll investigate how far self-directed learning can be extended, implementing novel approaches and shaping research directions.
Key responsibilities include:
- Training and evaluating models on proprietary RL environments to validate data quality, identify task coverage gaps, and close the feedback loop between environment design and model capability
- Architecting and optimizing RL training infrastructure, from training abstractions to distributed experiment management, using frameworks like Verl, OpenRLHF, or similar
- Designing, implementing, and testing training environments, evaluations, and methodologies for RL agents
- Profiling and optimizing end-to-end training runs to maximize experiment throughput and shorten research iteration cycles
The role offers competitive cash and equity compensation (>90th percentile), ownership and autonomy in a fast-moving startup environment, opportunity to collaborate with top machine learning engineers, comprehensive benefits including health/vision/dental, 401K match, daily onsite lunch, weekly snacks, and visa sponsorship with relocation support.
REQUIREMENTS:
- Experience running end-to-end LLM post-training pipelines with models at least 7B in size
- Proficiency in Python and PyTorch or JAX
- Experience with at least one modern RL training framework
- Experience building and operating ML infrastructure at scale
ADDITIONAL QUALIFICATIONS (preferred):
- Experience evaluating model outputs and building reward or evaluation signals
- Ability to stay current on post-training research and translate papers into running code
- Strong opinions on structuring RL training code for reproducibility and fast iteration
- Ability to balance research exploration with engineering rigor
- Strong systems design and communication skills
Note: PhD or extensive publications not required. The company values adaptability, exceptional communication, and collaboration skills.