SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 180,600 - 225,750 / annual
Scale AI is seeking a Machine Learning Research Scientist to join the Evaluations pod within the GenAI Research Organization. This role focuses on building benchmarks and diagnostic methods for frontier large language models and multimodal systems.
Key responsibilities include:
- Analyzing model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and AI agents, with emphasis on root cause analysis
- Designing and building benchmarks and evaluation methods that measure LLM capabilities across text and multimodal modalities
- Applying post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to data and training interventions
- Publishing research findings in top-tier AI conferences
- Collaborating with researchers and engineers to define best practices in evaluation-driven AI development
- Partnering with leading foundation model labs to translate failure analysis into technical and strategic input for next-generation generative AI models
Ideal candidates will have:
- Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or related field
- Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning
- Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning
- Track record of LLM evaluation or benchmark development
- Published research at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.)
- Excellent written and verbal communication skills
- Prior customer-facing research experience preferred
Scale AI works with industry-leading AI labs to provide high-quality data and accelerate GenAI research progress, partnering with organizations like Meta, Mayo Clinic, and U.S. government agencies.