SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's Critical Harm Operations team is seeking a senior cybersecurity practitioner and operations strategist to lead the Cyber vertical within User Safety & Risk Operations. This senior individual contributor role focuses on building enforcement systems for frontier risk and material harm that are accurate, fast, defensible, and scalable.
You will drive the cyber operations operating model across domain priorities, standard operating procedures, escalation paths, quality health, vendor capability, and roadmap inputs. You'll serve as the senior cyber expert for complex or high-risk decisions across ChatGPT, API, Codex, agents, and emerging product surfaces. The role requires translating policy ambiguity, quality gaps, appeals, and reviewer disagreement into clear decision rules, calibration examples, training, and tooling requirements.
Key responsibilities include building durable operating systems and quality loops (golden sets, holdouts, double-labeling, adjudication, error taxonomies, reviewer calibration, automation evaluations), raising FTE and BPO capability through onboarding and vendor partnerships, diagnosing root causes using quality and appeals signals, and building hands-on solutions using SQL, Python, dashboards, and LLM eval workflows.
You'll partner across Policy, Integrity, Safety Systems, Security, Legal, Product, Engineering, and Investigations to operationalize changes and drive launch readiness. Success is measured by durable improvement in the operating model and reviewer capability, not case volume.
Required: 8+ years hands-on cybersecurity experience in offensive security, threat intelligence, incident response, security research, red teaming, application security, DFIR, malware analysis, or related fields. Deep reasoning about attacker tradecraft, vulnerability exploitation, credential abuse, malware, persistence, evasion, and dual-use activity. Experience building or improving high-stakes operations, reviewer programs, QA systems, and vendor programs. Ability to translate cyber judgment into reviewer-usable SOPs and training. Comfort with SQL, Python, C/C++, JavaScript, PowerShell, Bash, APIs, and LLM tooling. Understanding of human-in-the-loop automation, evaluations, and monitoring. Experience with trust and safety, platform abuse, or LLM safety is valued.