SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Spectrum Labs (operating as Alice) is seeking a Research Lead for Evaluations and Benchmarks to own the design, execution, and release of AI safety and security benchmarks. You will ship a new benchmark every two to three weeks, each measuring frontier risks in AI systems that have not been previously evaluated. Some benchmarks are released publicly; others go only to leading AI labs. You collaborate with in-house researchers, leading AI labs, and universities on this work.
Key responsibilities include:
1. Benchmark Cadence: Ship benchmarks on a two to three week cycle. Scope varies by subject—chat-based taxonomies can include ~100 evals, while agentic or GRPO benchmarks are closer to 20 due to evaluation cost. You decide what ships publicly versus what goes to labs only based on sensitivity.
2. Quality Assurance: Set and maintain the quality bar. Frontier labs will rerun your benchmarks and verify they reproduce your numbers. You ensure verifiers hold, rubrics are clear, distributions are sound, and subject matter experts validate novelty. This means hands-on review of evals—you screen with models but also manually inspect items to catch taxonomy mismatches and push researchers for corrections.
3. Process Ownership: Hold the plan and calendar. You keep other researchers on timeline and can direct two to three freelance subject matter experts ad-hoc as needed.
4. Roadmap Setting: Monthly meetings with the CTO, pod leads, and research leads to set the quarterly release plan. Inputs include observations from research teams, client requests, and ecosystem developments. Output is a revised roadmap tied to business priorities.
5. Ecosystem Engagement: Spend ~20% of time staying ahead of the curve. Read emerging research, maintain relationships with lab researchers, attend conferences (2–3 per year), and have weekly conversations with lab contacts about emerging concerns.
Alice is a trust, safety, and security company serving the top 8 AI labs globally. The company provides end-to-end AI safety coverage: model hardening evaluations, pre-deployment red-teaming, runtime guardrails, and drift detection. You will work in the CTO office alongside the research lead who sets the public research agenda. Around 150 researchers work on harms directly across the organization, and you can pull any of them onto subjects as needed.
Requirements:
- PhD or Master's in computer science, machine learning, or related field, or equivalent depth from industry research
- 3+ years building and running safety or security evaluations for language models in production (at an AI lab, model provider, or safety/security research organization)
- 5+ relevant research publications in AI safety and security, with lead authorship on at least 2
- Strong engineering skills: evaluation harnesses, distributed inference, vLLM, code reading and debugging. Ability to build taxonomies, not just score against them
- Ability to direct researchers and freelancers without formal management authority
- Fluent English (written and spoken) for cross-timezone communication
- Curiosity about AI harms and willingness to learn new subjects every three weeks
- Ideal: post-training experience (SFT, DPO, GRPO); reward design for subjective/safety targets; agentic evaluation experience (tool use, orchestration, permissions, prompt injection); publications at top conferences; willingness to present work on client calls; strong presentation skills for senior audiences; availability to travel to conferences 3+ times per year