SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mercor is an AI data company on a mission to organize human intelligence to power the AI economy. The company operates a platform where millions of domain experts train frontier AI models, and has built the APEX benchmark family to measure AI's real-world impact on professional work. Mercor is a profitable Series C company valued at $10 billion.
As Research Scientist for APEX Benchmarks, you will lead the design of the next generation of APEX benchmarks and the expert-built datasets behind them. APEX is the AI Productivity Index—a family of benchmarks measuring whether frontier models can perform economically valuable professional work. APEX-1 covers single-turn tasks across investment banking, corporate law, consulting, and medicine. APEX-Agents tests multi-hour, cross-application agentic work in real tools. APEX-Accounting and APEX-SWE extend into accounting and real-world software engineering. Every task is written and graded by practicing experts on the Mercor platform, with results published as papers, open datasets, and public leaderboards that frontier labs monitor.
Key responsibilities include: deciding what the next APEX benchmark should measure based on frontier model performance and gaps in economically valuable work; owning task taxonomy, difficulty calibration, contamination controls, and statistical design; designing expert-built datasets and grading rubrics at scale; setting standards for result reporting including confidence intervals, inter-rater agreement, and failure analysis; partnering with academic collaborators and industry partners to co-design and adopt benchmarks; publishing research through arXiv papers, open datasets, blog posts, and conference talks; translating benchmark findings into clear ROI narratives for technical reports and customer conversations; and partnering cross-functionally with data operations, engineering, product, and strategy to move benchmarks from design to production.
You should have a strong applied or academic research background in LLM evaluation, benchmarking, NLP, or a related field with a track record of rigorous experimental design. You can identify informative measurements rather than easy-to-build ones, reason carefully about sampling, variance, contamination, and grader reliability, and possess strong coding skills to build eval harnesses and analyze results independently. Exceptional communication skills are essential—presenting complex technical findings clearly to frontier lab researchers and non-technical audiences. You thrive in fast-moving, cross-functional environments with ambiguous problem spaces and have genuine curiosity about GTM strategy, startup dynamics, and the business of AI data. This is a highly visible role at the intersection of research, company strategy, and go-to-market. You will work in-person five days a week in the San Francisco office in a high-intensity, high-ownership environment.
Nice-to-have qualifications include a Ph.D. in machine learning, NLP, or related field (or equivalent industry/frontier lab research experience); publications at top-tier venues (NeurIPS, ICML, ACL, ICLR) especially in evaluation, benchmarking, or data-centric AI; experience authoring a widely-adopted public benchmark or dataset; industry experience on evaluation, benchmarking, or post-training teams at frontier labs; domain depth in finance, law, consulting, accounting, medicine, or software engineering; and experience designing rubrics and model-as-judge pipelines.