SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's Future of Computing Research team is seeking a Research Scientist to advance multimodal AI systems that learn, adapt, and personalize over time. This role sits at the intersection of applied research and product development, focusing on RLHF (Reinforcement Learning from Human Feedback) and post-training methods for consumer-facing AI devices.
You will develop reward models and preference-learning pipelines that enable models to become more context-aware and adaptive. Key responsibilities include designing datasets, rubrics, and evaluation frameworks that capture user preferences and long-term value; running experiments on policy improvement using explicit feedback and implicit signals; and building long-horizon evaluation systems that measure whether model behavior actually improves outcomes over time.
The work is deeply grounded in real product use cases. You'll collaborate closely with safety researchers to ensure personalization remains aligned and interpretable, while working across the full stack—from data generation and labeling strategy through training runs, reward functions, and analysis. Success metrics go beyond benchmark performance to include trust, appropriateness, and demonstrable long-term user benefit.
Ideal candidates have strong backgrounds in machine learning research with experience in RLHF, reward modeling, preference optimization, or post-training for large models. You should have hands-on experience with reinforcement learning, ranking, recommender systems, personalization, or human-in-the-loop evaluation. You care about rigorous empirical work, can design clean experiments and reliable evaluations, and are comfortable with the challenge of training models against nuanced behavioral objectives. Experience building datasets or evaluation pipelines grounded in human preferences and real-world product behavior is valuable. You're excited by multimodal AI and how models can learn from richer interaction signals, and you thrive in close collaboration with engineers, designers, and safety researchers.
The role is based in San Francisco with a hybrid work model (3 days in office per week). Relocation assistance is available.