SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Anthropic is seeking a Staff Site Reliability Engineer to join the Safeguards ML Infrastructure team, which designs, builds, and operates the production infrastructure powering Claude's safety systems. This role sits at the center of operational work ensuring safeguards are properly configured and deployed for every frontier model release.
You will own critical responsibilities including: serving as launch captain for model releases, configuring and verifying safeguards across all deployment platforms (1P, AWS Bedrock, GCP Vertex); managing off-cycle deployment of new safety classifiers with canary rollouts and post-deploy validation; verifying safeguards are live on correct models and detecting configuration drift; and automating manual processes by converting launch runbooks into tooling and one-off deploys into repeatable pipelines.
Additional responsibilities include building and maintaining a safeguards registry with full provenance tracking, participating in on-call rotations for service incidents and time-sensitive launches, and serving as the safeguards point of contact during release windows.
Ideal candidates have 8+ years of software engineering or SRE experience with deep expertise in production change management at scale—including deploy pipelines, config management systems, canary analysis, and rollout safety. You should have run high-stakes releases as a launch captain or incident commander, have meaningful on-call experience with incident response and postmortem-driven improvements, and demonstrate a track record of automating operational toil. Hands-on experience deploying and operating AWS and GCP at scale is required. Proficiency in Python is essential; Rust experience is a plus. Familiarity with LLM inference systems and transformer-based model operations is valuable but not required. ML research background is not necessary—you will learn domain specifics on the job. The team values production judgment, the ability to ship changes to critical systems safely, and a drive to automate away last quarter's manual work.