SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's Safety Systems team is seeking a Model Policy Manager to shape how the company understands and addresses real-world risks from model misalignment as AI systems become more autonomous and operate over longer horizons.
You will investigate how misaligned behavior emerges across extended trajectories—including when models persist toward wrong objectives, take unsafe shortcuts, lose track of instructions, exploit environmental weaknesses, or circumvent constraints—and translate these insights into behavioral policies, evaluations, monitoring, and safeguards. This role bridges alignment research with the practical challenges of training and deploying frontier models.
Key Responsibilities:
- Identify vulnerabilities that emerge as models interact with tools, data, and external systems, translating them into model- and system-level safeguards
- Develop threat models and empirical frameworks for understanding harmful outcomes from misaligned behavior
- Identify underlying behaviors and system conditions driving those outcomes
- Turn findings into policy frameworks, evaluation criteria, online measurement, and safeguards
- Develop human data campaigns and gold sets to ground measurement and evaluation of emerging behaviors and risks
- Partner with research, engineering, security, and product teams to shape model and system safety, balancing trade-offs between safety, utility, and business risk
- Inform deployment decisions, system cards, safeguards reports, and OpenAI's broader approach to agentic safety
- Build monitoring approaches that detect regressions and emerging risks after deployment
Workplace: Hybrid (3 days in office per week in San Francisco; optional work-from-home Thursdays and Fridays). Relocation support available.
Requirements:
- Strong background in AI agent safety, privacy, security, cybersecurity, or adjacent fields, with an adversarial mindset to investigate real-world harmful outcomes
- Demonstrated interest in AI alignment and strong understanding of technical drivers of misaligned model behavior
- Technical fluency to work directly with evaluation and training data, understand what data shows, and identify limitations, patterns, and opportunities for deeper investigation
- Comfort working hands-on with model data and evaluation results: inspecting examples, analyzing failure patterns, assessing data quality, and distinguishing policy failures from grader, model, or system failures
- Ability to use empirical evidence to develop and refine safety policies and safeguards
- Capacity to translate complex or ambiguous alignment risks into precise behavioral expectations and measurable evaluation criteria
- Ability to work effectively across research, engineering, security, product, and policy teams
- Clear communication about complex and uncertain technical risks
- Thrives in fast-paced, collaborative research environments where priorities shift as models, evidence, and risks change