SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Preference Model is building automated ML research engineering to advance frontier AI models. The company is tackling a key bottleneck: the lack of high-quality reinforcement learning training environments that reflect real-world complexity. The founding team includes veterans from Anthropic's data team who built infrastructure and datasets behind Claude.
As a Member of Technical Staff in the Capabilities org, you will design and build reinforcement learning environments and reward functions that teach frontier models to perform ML research and engineering tasks. This role blends research and engineering, requiring you to stay current with cutting-edge research, develop novel approaches, and implement them in production systems.
Key responsibilities:
- Design and build RL environments and reward functions that produce clean, learnable signals for frontier models on ML research and engineering tasks
- Develop deep expertise across the frontier of ML research, training, and inference infrastructure
- Collaborate with researchers and engineers to brainstorm and create new tools to improve the environment-building process
- Own your environments end-to-end, from design through production deployment
- Conduct experiments and evaluations to validate your work
You will join a small, high-ownership team in a fast-moving startup environment, contributing directly to the data layer powering frontier LLM capabilities. The role offers full autonomy and the opportunity to work alongside top machine learning engineers.
QUALIFICATIONS:
- 5+ years of experience in machine learning or research, primarily on LLMs and transformer models
- Strong ML fundamentals and broad research interests; ability to read papers deeply and translate research into RL/VR problems
- Proficiency in Python and systems programming; expertise in at least one of PyTorch or JAX
- Problem-solving mindset with ability to drive solutions end-to-end
- Passion for staying current with rapidly evolving ML infrastructure
- Ability to meet throughput expectations and respond quickly to feedback
PLUS (desired):
- Expert knowledge in an active deep learning/ML research area with publications or public code; research background (PhD, MS) is a strong plus
- Deep understanding of transformer internals and training/inference of modern LLMs; experience with inference libraries (vLLM, SGLang, etc.)
- Strong expertise in kernel development (CUDA, Triton, Pallas)
- Experience building complex interactive RL environments