SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 500,000 - 850,000 / annual
Anthropic is seeking a Staff Research Engineer to join the Multi-Agent Scaling team. This role sits at the intersection of research and engineering, focusing on how large teams of agents solve problems that individual agents cannot. The team studies performance, cost, and coordination as agent count, compute budget, and task length scale, building the platform and evaluations Anthropic uses to run and measure large agent teams.
You will design and run large-scale experiments on agent teams, investigating how performance and efficiency change with team size, compute, and task horizon. You'll build and scale systems that reliably run very large agent teams, debug failures that only appear at scale, and design evaluations for long-horizon problems while maintaining result trustworthiness. Additional responsibilities include building tooling and metrics for researchers to understand agent team behavior, and partnering across Anthropic to enable other research teams to run experiments on the platform.
Representative projects include measuring performance scaling with agent count, preparing large-scale agent runs by identifying and fixing bottlenecks, optimizing compute budget allocation across agent teams, building visualization tooling for large agent team activity, and designing novel evaluations that distinguish real teamwork gains from evaluation artifacts.
This is a generalist role on a small team where you'll move quickly from vague questions to running experiments, combining research rigor with engineering pragmatism.
REQUIREMENTS:
- Significant software engineering, ML, or research engineering experience
- Demonstrated ownership of something substantial end-to-end (large system, evaluation/benchmark, agent product, or research project)
- Genuine enjoyment of both research and engineering work
- Quantitative thinking about complex systems with healthy skepticism of numbers
- Ability to work from vague questions rather than detailed specifications
- Results-oriented with bias toward flexibility and impact
- Clear written and verbal communication
- Care about societal impacts of AI work
- Bachelor's degree or equivalent combination of education, training, and professional experience in a field relevant to the role
STRONG CANDIDATES MAY ALSO HAVE:
- Experience building or operating large-scale distributed systems (schedulers, sandboxed code execution, inference/RL infrastructure)
- Built evaluations, benchmarks, or harnesses for LLMs or agents
- Built complex agentic systems using LLMs
- Experience with scaling laws or large-scale empirical research
- Background in operations research, statistics, economics, physics, quantitative finance, or similar fields that model and optimize complex systems
NOTE: Formal certifications, academic research experience, publication history, and prior multi-agent systems or reinforcement learning experience are not required.