SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's Interpretability team is seeking a researcher to study internal representations of deep learning models and develop techniques for understanding how AI systems work. The role focuses on mechanistic interpretability—reverse-engineering neural networks to understand their behavior—with direct application to AI safety and alignment.
You will develop and publish research on interpretability techniques, engineer infrastructure for studying model internals at scale, and collaborate across OpenAI teams on projects leveraging the company's unique computational resources. The work directly supports OpenAI's mission to ensure advanced AI systems remain safe and aligned as they grow more capable.
Key responsibilities include:
- Designing and executing research programs in mechanistic interpretability
- Building tools and infrastructure for analyzing model representations at scale
- Publishing findings and contributing to the broader AI safety research community
- Guiding research directions toward practical impact and long-term scalability
- Cross-team collaboration on safety-critical interpretability projects
Ideal candidates bring a Ph.D. or equivalent research experience in computer science, machine learning, or related fields, with 2+ years of research engineering experience. You should have proficiency in Python or similar languages, deep familiarity with large-scale AI systems, and genuine enthusiasm for AI safety and alignment. Prior work in mechanistic interpretability, AI safety, or related areas is valued. Above all, you should be deeply curious about how neural networks work and motivated by the challenge of making powerful AI systems more transparent and controllable.