SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's Safety Systems team is seeking a Model Policy Manager to define how frontier AI models should behave in high-risk and high-ambiguity contexts, including agentic systems, multimodal systems, user safety, privacy, and emerging risk domains.
You will design and maintain model policies across safety-relevant domains, translating risk and harm models into clear behavioral specifications, evaluation criteria, and system-level safeguards. Key responsibilities include:
- Design and maintain model policies across dual-use, agentic, and emerging frontier-risk areas
- Translate risk models into behavioral specifications, evaluation criteria, and grading guidance
- Define practical boundaries between beneficial AI uses and assistance that could enable harm or misuse
- Build policy artifacts supporting model training, evaluation, and deployment
- Partner with safety researchers, engineers, product teams, and stakeholders to operationalize policy into scalable model behavior
- Use red-teaming results, deployment data, and model failures to improve policy and evaluation quality
- Identify emerging capability areas where frontier AI could create new safety challenges
- Study real-world deployments to identify where model behavior succeeds, fails, or drifts from intended safety posture
- Combine longer-horizon safety research with hands-on launch and deployment work
- Contribute to system cards, safety reports, policy documentation, and external communications
- Design and run human data campaigns including gold set construction, labeling guidance, calibration, and eval coverage analysis
This is a hybrid role based in San Francisco (three days in office per week, optional work from home Thursdays and Fridays). OpenAI offers relocation support.
Requirements:
- Strong judgment about how advanced AI systems affect real-world risk, especially in ambiguous, fast-moving, or high-impact areas
- Experience building or applying policies, taxonomies, harm models, threat models, or risk frameworks for complex technical, social, or adversarial systems
- Ability to move across domains without being the deepest subject-matter expert in every area, while knowing when to seek expert input
- Demonstrated ability to turn fuzzy questions into structured policy frameworks, evaluation criteria, operational guidance, and enforceable model behavior
- Comfort using empirical evidence (evaluations, red-teaming results, deployment observations, model failure modes) to inform policy decisions
- Systems thinking across policy, data, graders, classifiers, training, deployment safeguards, measurement, monitoring, and escalation workflows
- Technical judgment about what model behavior can realistically be trained, measured, evaluated, and enforced at scale
- Strong cross-functional collaboration skills with research, engineering, product, policy, domain experts, and operational teams
- Clear writing ability about complex tradeoffs involving safety, user value, and implementation constraints
- Pragmatic approach to safety focused on reducing real-world risk while preserving legitimate and beneficial AI uses
- Comfort in fast-paced, collaborative research environments where priorities shift with models, evidence, and risks
- Grounding in implementation details, empirical results, and what can actually be trained or measured