SlipstreamJobsFresh Startup & VC-Backed Jobs

Machine Learning Research Scientist, Evaluations

Scale AI - Seattle, WA, United States - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 180,600 - 225,750 / annual

Scale AI is seeking a Machine Learning Research Scientist to join the Evaluations pod within the GenAI Research Organization. This role focuses on building benchmarks and diagnosing failure modes in frontier large language models and multimodal systems. You will develop rigorous evaluation frameworks and diagnostic methods that identify where state-of-the-art models fail and why. Key responsibilities include analyzing model behavior to characterize failure modes across capability gaps, reasoning errors, robustness issues, and alignment concerns. You'll design and build benchmarks that measure LLM capabilities in both text and multimodal modalities, applying post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to specific data and training interventions. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development and partner with leading foundation model labs to translate failure analysis into technical and strategic guidance for next-generation generative AI systems. Published research findings in top-tier AI conferences is expected. Ideal candidates hold a Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or related field with deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. You should have hands-on experience with post-training techniques (RLHF, preference modeling, instruction tuning) and LLM evaluation or benchmark development. Strong written and verbal communication skills are essential, along with a track record of published research at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.). Prior customer-facing experience is valued. Scale works with industry-leading AI labs including Meta, providing high-quality data and full-stack technologies that power the world's most important AI systems. The company also partners with enterprises and government agencies to build, deploy, and oversee AI applications.

Similar roles