SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mercor is an AI data company on a mission to organize human intelligence to power the AI economy. The company operates a platform where millions of domain experts train frontier AI models, and has built the APEX benchmark family to measure AI's real-world impact on professional work. Mercor is a profitable Series C company valued at $10 billion, with offices in San Francisco, NYC, and London.
As a Research Engineer in Benchmarking, you will work at the intersection of engineering and applied AI research, owning benchmarking pipelines, evaluation systems, and failure analysis workflows that directly inform how frontier language models are trained and improved.
Key responsibilities include:
- Designing, implementing, and maintaining benchmarks and metrics for tool use, agentic behavior, and real-world reasoning, ensuring they scale with training and align with product and research goals
- Building and operating end-to-end LLM evaluation systems including runs, scoring, dashboards, and reporting so researchers can track model performance and compare runs at scale
- Running systematic failure analysis on model outputs (wrong tool use, reasoning errors, safety/alignment issues), categorizing failure modes, quantifying prevalence, and feeding findings into reward design and data curation
- Creating and refining rubrics, automated evaluators, and scoring frameworks that drive training and evaluation decisions, balancing rigor with scalability
- Quantifying data usability, quality, and impact on key benchmarks; using evals and failure analysis to guide data generation and curation
- Collaborating with AI researchers, applied AI teams, and data producers to align evals with training objectives and prioritize benchmarks that matter most
- Operating with strong ownership in a high-iteration research setting
Required qualifications include strong applied research background in model evaluation, benchmarking, or failure analysis; strong coding skills and hands-on ML experience; solid grasp of data structures, algorithms, and backend systems; comfort with APIs, SQL/NoSQL, and cloud platforms; and ability to reason about model behavior and data quality. The role requires in-person work five days a week in San Francisco in a high-intensity, high-ownership environment.
Nice-to-have qualifications include industry experience on post-training or evaluation/benchmarking teams, publications at top-tier venues (NeurIPS, ICML, ACL), experience building LLM evaluations or benchmarks, experience with synthetic data generation or RL-style workflows, and work samples demonstrating relevant skills.