SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 320,000 - 485,000 / annual
Anthropic is seeking a Staff Software Engineer to join the Safeguards team, responsible for building safety and oversight mechanisms for AI systems. This role focuses on developing systems to detect unwanted model behaviors, prevent misuse, and ensure user well-being through technical implementation of safety, transparency, and oversight principles.
Key responsibilities include:
- Developing monitoring systems to detect unwanted behaviors from API partners with automated enforcement actions and internal dashboards for analyst review
- Building abuse detection mechanisms and infrastructure at scale
- Surfacing abuse patterns to research teams to harden models during training
- Constructing robust, multi-layered real-time defenses for safety mechanisms that operate at production scale
You will work across the full stack, applying technical expertise to enforce terms of service and acceptable use policies while collaborating with operational teams and research groups.
Anthropicis a public benefit corporation headquartered in San Francisco, focused on creating reliable, interpretable, and steerable AI systems. The team is collaborative and values high-impact research with emphasis on communication skills and cross-functional collaboration.
Current hybrid policy requires staff to be in office at least 25% of the time, though some roles may require more. Visa sponsorship is available.
REQUIREMENTS:
- Bachelor's degree in Computer Science, Software Engineering, or comparable experience
- Proficiency in Python and TypeScript
- Ability to work across the stack
- Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders
STRONG CANDIDATES MAY ALSO HAVE:
- 8+ years of software engineering experience
- Experience with integrity, spam, fraud, or abuse detection and mitigation
- Experience building trust and safety detection mechanisms for AI/ML systems
- Experience with prompt engineering, jailbreak attacks, and adversarial inputs
- Experience building custom internal tooling with operational teams