SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Protege is building a platform to solve AI's biggest bottleneck: access to high-quality training data. The company facilitates secure, efficient, and privacy-centric exchange of AI training data and is backed by world-class investors with partnerships across leading AI teams.
You will be the first Machine Learning Engineer dedicated to the Benchmarks and Evaluations vertical, working directly with the General Manager and research team to establish Protege as a market leader in AI evaluation infrastructure.
Key responsibilities include:
**Eval Foundation**: Partner with the GM and early customers to define strong evaluation standards across different domains. Work with Protege researchers to design and build benchmarks, establishing standards for processing different data modalities.
**Infrastructure Ownership**: Build the backend infrastructure for the vertical, including data pipelines, execution environments, storage, and orchestration systems. Stand up sandboxed environments for agentic evaluations where models require tools, code execution, or multi-step task capabilities.
**Product Development**: Identify repeatable evaluation patterns, infrastructure gaps, and product opportunities from live customer engagements. Collaborate with the DataLab research team on domain-specific data and research questions.
In your first 90 days, you'll build understanding of the evaluation landscape, the vertical's strategy, and customer demand. You'll identify the largest technical bets and ship multiple iterations of eval infrastructure while owning the engineering portion of customer engagements end-to-end.
Required qualifications: 4+ years of engineering experience with hands-on ML work evaluating models. You must have previously owned backend and infrastructure systems, demonstrate high ambiguity tolerance with bias toward action, and communicate effectively in writing.
Preferred experience includes building benchmarks or evals for LLMs, time at frontier labs or research organizations, founding/early-stage startup experience, and familiarity with agentic systems, RL environments, code-execution sandboxes, and trusted execution environments.