SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's Safety Training research team is seeking a researcher to advance capabilities for implementing safe behavior in AI models, with a focus on U.S. government and national security applications. The role bridges research, engineering, security, and policy to ensure deployed models are safe, beneficial, and trustworthy in safety-critical situations.
In this role, you will research and implement methods for safety training, reinforcement learning, and adversarial robustness. You'll develop evaluations to identify model failure modes and use findings to improve training approaches. You'll collaborate with research, engineering, security, and policy partners to support safe and reliable deployment of models in national security contexts.
The team's focus areas include training nuanced safety behaviors, making models robust to adversarial actors, addressing privacy and security risks, and ensuring trustworthiness in safety-critical deployments. You'll help advance post-training safety and robustness while preserving model usefulness and capabilities.
REQUIREMENTS:
- 4+ years of relevant AI safety research experience, including RLHF, adversarial training, or robustness
- Degree in computer science, machine learning, or related field
- Strong deep learning research or engineering skills
- Experience improving model safety for deployment
- Active TS/SCI clearance or equivalent (security requirement)
- Collaborative research mindset and motivation aligned with responsible AI deployment