SlipstreamJobsFresh Startup & VC-Backed Jobs

Safeguards Enforcement Analyst, User Well-being

Anthropic - San Francisco, CA, United States - Hybrid - posted 2026-08-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Anthropic is seeking a Safeguards Enforcement Analyst to join the User Well-being team, focused on supporting the design and deployment of mental health guardrails for Claude. This role involves iterating on detection systems, managing review queues, evaluating new interventions, and monitoring existing safeguards. The team addresses interconnected harms including suicide, self-harm, disordered eating, AI sycophancy, and emotional dependence on AI. Key responsibilities include supporting the design and execution of interventions with defined metrics and evaluation datasets; partnering with Engineering and Data Science teams to build, tune, and validate detection models for automated systems, including threshold-setting and precision/recall tradeoffs; monitoring intervention and detection system performance over time; reviewing flagged content to drive enforcement and policy improvements; supporting in-product features that connect users to crisis resources in collaboration with Product, Legal, and external partners; providing detailed feedback to the Safeguards Policy Design team based on real scenarios; and staying current with emerging AI policy and external research on AI's relationship to mental health. Minimum qualifications include experience in trust & safety, product policy, content moderation, or related fields with direct exposure to mental health, suicide, self-harm, or related well-being harms; experience designing or running experiments, evaluations, or measurement studies; experience translating policy definitions into measurable form (rubrics, review guidelines, classification criteria); experience managing or coordinating content review operations including quality assurance and workflow management; proficiency in SQL and/or other data analysis tools; experience working with generative AI products including prompt engineering for content review and classification; ability to turn open questions and data into concise analysis; experience identifying emerging risks and communicating findings to cross-functional stakeholders; understanding of challenges in implementing product policies at scale in content moderation; and sound judgment in ambiguous, high-consequence cases. Preferred qualifications include subject matter expertise in mental health from academia, clinical practice, crisis intervention, or trust & safety; experience building or evaluating LLM-based classification systems; experience using agentic tools to scale analysis or automate work; and experience working within crisis support. Note: This position involves exposure to explicit content spanning sexual, violent, and psychologically disturbing topics.

Similar roles