SlipstreamJobsFresh Startup & VC-Backed Jobs

Research Manager, Biological Safety

Anthropic - San Francisco, CA, United States - In-office - posted 2026-09-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Anthropic is hiring a Research Manager to lead the Biological Safety team within the Safeguards organization. This is a hands-on management role focused on building policies, evaluations, and enforcement systems that prevent AI models from contributing to catastrophic harm, specifically in the biological domain. You will manage a team of research scientists and engineers responsible for designing capability evaluations against frontier models, curating training data for safety classifiers, training and iterating on classifiers with ML engineers, and measuring performance against adversarial pressure in production. You'll set technical direction, decide team investments, and own results. Key responsibilities include: managing and growing the team through hiring, onboarding, and career development; setting the technical roadmap for biological safety research; owning the quality of capability evaluations and translating results into deployment recommendations; guiding development of training and evaluation datasets with threat modeling experts; overseeing classifier training and iteration to optimize for adversarial robustness and low false-positive rates; ensuring investment in tooling and pipelines for fast, repeatable evaluation; establishing performance measurement against production traffic; directing red-teaming and stress-testing; partnering with Research, Product, Policy, and government affairs teams; representing work in external communications including model cards and policy documents; and tracking developments in biology, ML, and biosecurity. The core tension the team owns is precision: safeguards must be robust against sophisticated actors while remaining transparent to the far larger population of legitimate researchers using Claude for life sciences work. This is an empirical problem requiring careful measurement and tradeoff analysis. You'll maintain enough technical depth to review evaluation designs, interrogate classifier failure modes, and represent the work credibly to partners across the organization.

Similar roles