SlipstreamJobsFresh Startup & VC-Backed Jobs

Model Policy Manager

OpenAI - San Francisco, CA, United States - Hybrid - posted 2026-09-16

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

OpenAI's Safety Systems team is seeking a Model Policy Manager to define how frontier AI models should behave in high-risk and high-ambiguity contexts, including agentic systems, multimodal systems, user safety, privacy, and emerging risk domains. You will design and maintain model policies across safety-relevant domains, translating risk and harm models into clear behavioral specifications, evaluation criteria, and system-level safeguards. Key responsibilities include: - Design and maintain model policies across dual-use, agentic, and emerging frontier-risk areas - Translate risk models into behavioral specifications, evaluation criteria, and grading guidance - Define practical boundaries between beneficial AI uses and assistance that could enable harm or misuse - Build policy artifacts supporting model training, evaluation, and deployment - Partner with safety researchers, engineers, product teams, and stakeholders to operationalize policy into scalable model behavior - Use red-teaming results, deployment data, and model failures to improve policy and evaluation quality - Identify emerging capability areas where frontier AI could create new safety challenges - Study real-world deployments to identify where model behavior succeeds, fails, or drifts from intended safety posture - Combine longer-horizon safety research with hands-on launch and deployment work - Contribute to system cards, safety reports, policy documentation, and external communications - Design and run human data campaigns including gold set construction, labeling guidance, calibration, and eval coverage analysis This is a hybrid role based in San Francisco (three days in office per week, optional work from home Thursdays and Fridays). OpenAI offers relocation support. Requirements: - Strong judgment about how advanced AI systems affect real-world risk, especially in ambiguous, fast-moving, or high-impact areas - Experience building or applying policies, taxonomies, harm models, threat models, or risk frameworks for complex technical, social, or adversarial systems - Ability to move across domains without being the deepest subject-matter expert in every area, while knowing when to seek expert input - Demonstrated ability to turn fuzzy questions into structured policy frameworks, evaluation criteria, operational guidance, and enforceable model behavior - Comfort using empirical evidence (evaluations, red-teaming results, deployment observations, model failure modes) to inform policy decisions - Systems thinking across policy, data, graders, classifiers, training, deployment safeguards, measurement, monitoring, and escalation workflows - Technical judgment about what model behavior can realistically be trained, measured, evaluated, and enforced at scale - Strong cross-functional collaboration skills with research, engineering, product, policy, domain experts, and operational teams - Clear writing ability about complex tradeoffs involving safety, user value, and implementation constraints - Pragmatic approach to safety focused on reducing real-world risk while preserving legitimate and beneficial AI uses - Comfort in fast-paced, collaborative research environments where priorities shift with models, evidence, and risks - Grounding in implementation details, empirical results, and what can actually be trained or measured

Similar roles