SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
HUD is building infrastructure for RL training data and evals for frontier AI agents, with a marketplace connecting data creators to frontier labs. The company has raised $16M from top VCs and was part of Y Combinator W25.
As a Software Engineer, Data Quality, you will build full-stack tools that help HUD's Quality and Intelligence (QNI) team assess and improve training data and evals for frontier agents. You'll transform hands-on quality workflows into reliable internal products and bring the most useful capabilities into HUD's production platform.
Key responsibilities include:
- Building full-stack tools for reviewing tasks, inspecting agent trajectories, checking graders, and investigating data quality issues
- Creating dashboards and metrics that show quality trends, failure modes, and the health of review and validation workflows
- Designing backend services, APIs, and data workflows that connect quality checks with task creation, evaluation, and feedback to data creators
- Working with QNI and research engineers to turn evolving review methods into clear, efficient workflows that scale beyond manual analysis
- Moving proven internal tools into production, improving their reliability, usability, and observability as adoption grows
- Investigating issues in live workflows and using insights to improve the platform and prevent repeat failures
You'll work closely with QNI and research engineers across task review, trajectory inspection, grader checks, quality metrics, and feedback mechanisms. The work spans interfaces, backend services, and data workflows, with a focus on making quality issues easier to find, understand, and fix.
The team currently has ~25 people, mostly full-time in-person but with some remote flexibility. The team includes 4 International Olympiad medalists, serial AI startup founders, and researchers with publications at top venues like ICLR and NeurIPS. The company is scaling profitably with strong demand and has 8 figures in funding.
The company offers competitive compensation, 100% covered top-of-the-line medical, dental, and vision insurance (US employees), lunch and dinner in office, company-wide holiday break, Equinox membership, 401k, commuter benefits (US employees), and unlimited access to tokens for ChatGPT, Claude Code, Cursor, and similar tools.
REQUIREMENTS:
You should have:
- Strong software engineering fundamentals and the ability to build across frontend, backend, and data systems
- Proficiency in Python and a modern web stack such as TypeScript and React, or comparable tools
- Experience building internal or user-facing products end-to-end, from understanding a workflow through shipping and improving it
- Sound judgment about APIs, data models, and production systems, including how to make them reliable and easy to debug
- Ability to turn ambiguous quality problems into useful interfaces, automation, and measurable checks
- Clear communication and comfort working closely with research and quality teams to understand how people use the tools you build
Strong candidates may also have:
- Built tooling for data review, annotation, evals, benchmarks, or ML workflows
- Worked with agent traces, graders, reward signals, or RL training data
- Designed dashboards, observability tools, or review workflows for complex datasets and pipelines
- Improved a prototype or internal tool until it was ready for wider production use
The company prioritizes technical aptitude and learning potential over years of experience. Visa sponsorship is provided for relocation to the US or Singapore. The hiring process includes 2 technical interviews and a 2-3 day work trial.