SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
HUD is building infrastructure for reinforcement learning training data and evaluations for frontier AI agents, with a marketplace connecting these capabilities to leading AI labs. The company has raised $16M from top VCs, was part of Y Combinator W25, and serves frontier labs, Fortune 500 companies, and startups.
As Research Manager, you will lead a team of research engineers on projects that make HUD's agent training data and evaluations more effective for improving frontier models. You'll work at the intersection of technical depth and team leadership, staying engaged in the research while guiding the team through ambiguous problems, rigorous experimentation, and scalable implementation.
Key responsibilities include:
- Setting research direction for data quality, defining how HUD measures whether tasks, trajectories, rewards, and evaluations are reliable and useful for agent training
- Leading research engineers from problem definition through experiments, implementation, and clear conclusions; mentoring them to develop stronger technical judgment
- Designing and reviewing experiments that connect model behavior and failure modes to data, environment, and reward design
- Developing methods for validating and improving training data at scale, including trajectory audits, grader checks, and feedback loops
- Partnering with research engineers, domain experts, and data vendors to translate research insights into workflows, tools, and quality standards
- Communicating findings and tradeoffs clearly to inform team prioritization and cross-functional learning
The team currently has ~25 people, mostly full-time in-person but with some remote flexibility. The team includes 4 International Olympiad medalists, serial AI startup founders, and researchers with publications at top venues (ICLR, NeurIPS, etc.). The company is scaling profitably with strong demand.
Qualifications:
- Experience leading technical research projects to completion, from open question through evidence, decision, and working result
- Direct experience managing and mentoring researchers or research engineers while remaining engaged in technical work
- Strong understanding of machine learning and reinforcement learning, including how training objectives, data, and feedback shape model behavior
- Experience with agent training data, evaluations, benchmarks, synthetic data, or model evaluation infrastructure
- Sound experimental judgment to distinguish useful training signals from metrics that only appear convincing
- Strong written communication and ability to explain methods and findings to researchers, engineers, and external partners
Additional strengths valued:
- Experience building scalable data quality systems or validation pipelines for model training
- Experience diagnosing reward hacking, grader errors, or other subtle agent failure modes
- Track record translating research findings into tools and processes used by others
- Early-stage startup experience and strong cross-team collaboration skills
The company prioritizes technical aptitude and learning potential over years of experience.