SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 350,000 - 850,000 / annual
Anthropic is seeking a Research Scientist to join the Takeoff Intel team, focusing on measuring and understanding recursive self-improvement in large language models. This role combines hands-on technical research with strategic impact, requiring deep expertise in model development loops and the ability to identify which signals truly matter for AI R&D acceleration.
You will design and build evaluations that measure capability growth and self-improvement dynamics, grounded in real evaluation and telemetry data. Key responsibilities include identifying signals that track AI R&D acceleration, constructing quantitative models of capability growth, running experiments to test hypotheses about automation and capability, and making opinionated research bets while owning outcomes. You'll also write graded assessments of measurement results for internal decision-makers and public reporting, and collaborate across pretraining, RL, economic research, and policy teams.
Ideal candidates have hands-on research experience with large language models (pretraining, fine-tuning, RL, evals, or agent systems), strong quantitative instincts, and comfort with quantitative modeling. Experience in forecasting, AI capability assessment, or scaling laws is valuable. You should be able to design rigorous evaluations from vague questions, write clearly with calibrated confidence statements, and be motivated by impact—understanding that outputs may be graded assessments and system-card sections rather than traditional papers. A background in physics, applied math, or similar quantitative fields that transitioned to ML is a plus. Care about AI safety and the implications of rapid capability growth.
Anthropic's recent work includes the Epoch Capabilities Index (ECI) adapted for system cards, AI R&D capability assessments in Claude system cards, and the "When AI Builds Itself" research article. The team operates at the frontier of AI development and safety.