SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mercor is a Series C AI data company ($10B valuation) building infrastructure between human expertise and frontier AI models. The company operates a platform where millions of domain experts train AI models, and has developed APEX, an AI Productivity Index that benchmarks how effectively frontier models perform economically valuable work.
As Product Manager for APEX, you will own the strategy, roadmap, and operational execution of Mercor's evaluation products, benchmarks, and public/private leaderboards. This is a high-impact role at the intersection of AI research, product, and go-to-market.
Key responsibilities include:
- Own the complete roadmap and portfolio strategy for benchmark development, leaderboard launches, and infrastructure investment. Decide which evaluation domains are prioritized, when benchmarks saturate, and what replaces them.
- Run intake and prioritization for new benchmark proposals from research, customers, and go-to-market teams against real demand and company strategy.
- Enforce evaluation integrity by owning contamination policy, holdout strategy, versioning, auditability, and release cadence. Publish methodology with sufficient clarity that external researchers can reconstruct results.
- Build the end-to-end evaluation pipeline from execution through publication, including harness execution, grading, model onboarding, hyperparameter scaffolds, cost/latency reporting, and leaderboard surfaces.
- Work directly with frontier AI labs and enterprise customers to understand evaluation needs, present and defend results, and translate feedback into the next generation of benchmarks. Support go-to-market on launches, partnerships, and thought leadership.
- Connect leaderboard demand signals to business outcomes. Track adoption, usage, and downstream revenue to justify ROI and inform loss analysis investments.
- Operate hands-on: write specs and PRs, spot-check failures, and unblock yourself to continuously improve systems.
You'll partner with world-class benchmark researchers (first authors of Tau Bench, SciCode, PostTrainBench, and others), engineering teams, and operations. The role requires balancing research rigor with product pragmatism, and representing Mercor as a thought leader in AI evaluation.
Mercor operates in-person five days per week from San Francisco, NYC, or London offices. The company is profitable and well-funded, offering competitive compensation including equity, performance bonuses, relocation support, and comprehensive benefits.