SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Handshake AI is a rapidly scaling AI data business (grown from $0 to ~$1B run rate in 2025) that partners with frontier AI labs to create evaluations, benchmarks, and training data. The company powers 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.
As Strategic Projects Lead for Safety & Red Teaming, you will own end-to-end execution of large-scale adversarial testing programs that evaluate frontier AI models before and after release. You will design and run structured efforts to push models to failure—including evaluating harm and policy categories, developing jailbreak and adversarial technique pipelines, and extending red-teaming into agentic settings—then work directly with frontier labs to translate findings into sharper policies and stronger evaluation coverage.
Key responsibilities:
- Own end-to-end execution of red-teaming and safety evaluation programs spanning sensitive, high-severity harm categories, from scoping through delivery
- Run multiple concurrent programs across different frontier labs and policy domains, each with distinct scope, harm categories, and stakeholders
- Design and execute adversarial testing methodologies (jailbreak technique development, iterative push-to-failure testing) to surface model vulnerabilities against frontier labs' policies
- Extend red-teaming methods into agentic contexts, evaluating model behavior under adversarial pressure with tools and multi-step autonomy
- Lead and coordinate teams of expert Fellows executing sensitive, high-stakes testing work, maintaining precision and consistency across harm categories, languages, and markets
- Partner directly with policy leads at frontier AI labs to identify coverage gaps, ambiguous classifications, and emerging risk areas; help shape policy documents based on findings
- Synthesize testing data into insight reports surfacing trends, systemic gaps, and recommended areas of policy and evaluation development
- Design and adapt staffing models for your red-teaming workforce (team size, skill mix, training, incentive structures) to improve throughput and delivery reliability
This is a high-ownership, outcomes-driven role for operators who thrive in ambiguity, move fast with incomplete information, and are accountable for results at scale. You will make real-time decisions affecting delivery quality, program scope, and long-term customer relationships with frontier AI labs.
Note: This role involves exposure to sensitive or explicit content (violent, sexual, or otherwise disturbing material) as part of safety and policy evaluation work.
REQUIREMENTS:
- 2+ years of experience in trust & safety, red-teaming, security research, policy enforcement, or a related technical/analytical field
- Excellent project and workforce management skills; comfortable leading and directing teams of red-teamers and testers against tight timelines and shifting scope
- Strong analytical and first-principles problem-solving skills; comfortable operating in ambiguous, fast-changing testing environments
- Working familiarity with adversarial testing concepts (jailbreak techniques, prompt-based exploits) or demonstrated ability to pick them up quickly
- Exceptional communication and stakeholder management skills, including with senior customers and policy teams at frontier AI labs
- High ownership mindset with pride in end-to-end accountability
- Curiosity and ability to quickly learn technical AI concepts, model behavior, and industry trends
BONUS:
- Experience red-teaming, jailbreaking, or adversarially evaluating LLMs or agentic AI systems
- Subject matter expertise in high-severity policy harms (e.g., CBRN, cyber, violent extremism)
- Experience running testing or evaluation programs
- Background in security research, offensive security, or AI safety research