SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Reflection AI is hiring a Member of Technical Staff for the Data Flywheel team, a hands-on technical role at the intersection of research and deployment. The Data Flywheel team closes the gap between benchmark performance and useful performance in the real world by identifying and building signals, data, and feedback loops that turn model usage into rigorous evaluations, targeted training data, and measurable improvements in future model generations.
In this role, you will identify high-value data sources and partnership opportunities, deeply understanding underlying use cases and translating them into representative evaluations. You'll bring new data sources online from initial partner conversations through data scoping, quality validation, and integration into production evaluation and training pipelines. You'll design and build evaluations, graders, and feedback loops that make priority real-world model behaviors measurable.
You'll analyze model performance and failure modes, translating insights into targeted datasets, reward signals, and training interventions. You'll develop human and synthetic data strategies for capabilities where existing data is insufficient, including designing and running evaluation and data-collection programs with vendors. You'll build infrastructure and pipelines needed to ingest, inspect, version, and evaluate data reliably at scale.
The role requires collaboration across pre-training, post-training, applied, and partnership teams to turn new signals into measurable model improvements. You'll work with researchers and engineers across the company, as well as customers, partners, vendors, and the open-source community.
Required qualifications include a degree (BS, MS, or PhD) in Computer Science, Machine Learning, or related discipline, or equivalent practical experience. You need deep technical understanding of LLM training and evaluation with hands-on experience in evaluation design, data curation, reinforcement learning, or reward design. Strong software engineering skills and experience building automated data/evaluation pipelines or large-scale ML systems are essential. You should have a track record of owning high-impact projects end-to-end, navigating ambiguity, and adapting quickly as priorities change. A highly collaborative, action-oriented approach with excitement about defining how a frontier lab measures and accelerates model progress is critical. High agency and thriving in a fast-paced startup environment with bias for impact over process are expected.