SlipstreamJobsFresh Startup & VC-Backed Jobs

Researcher, Agent Safety, Oversight and System Mitigations

OpenAI - San Francisco, CA, United States - Hybrid - posted 2026-09-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

OpenAI's Agent Safety team is seeking a researcher or engineer to focus on oversight and system-level mitigations that enable increasingly capable AI agents to operate safely and autonomously in real environments. The team's mission is to reduce the probability of severe unintended outcomes from advanced AI agents while preserving their ability to act effectively. The Agent Safety team works across three core areas: training methods and environments that teach agents to make better decisions in consequential situations; measurements and evaluations that identify emerging risks; and oversight mechanisms that reduce harmful actions while preserving useful autonomy. In this role, you will design, build, and evaluate system-level controls for agent actions, including agent-based review systems. You'll work closely with engineering teams to productionize AI controls and red-team end-to-end agentic systems to measure whether controls prevent data exfiltration, unsafe tool use, and other harmful outcomes. A key focus is improving the safety-productivity tradeoff by measuring and reducing missed harmful actions, unnecessary blocks, approval burden, and latency. You should have strong systems or security instincts and can reason concretely about isolation boundaries, permissions, attack surfaces, and failure modes in complex systems. The ideal candidate enjoys turning ambiguous safety questions into concrete threat models, reproducible experiments, and practical mitigations, and can revise approaches based on evidence from deployment. Experience building robust experimental infrastructure and designing evaluations that distinguish promising mitigations from brittle ones is valuable. A background in AI control or security is welcome but not required. This is a hybrid role based in San Francisco with 3 days in the office per week. OpenAI offers relocation assistance to new employees.

Similar roles