SlipstreamJobsFresh Startup & VC-Backed Jobs

Machine Learning Research Scientist, Evaluations

Scale AI - San Francisco, CA, United States - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 180,600 - 225,750 / annual

Scale AI is seeking a Machine Learning Research Scientist to join the Evaluations pod within the GenAI Research Organization. This role focuses on building benchmarks and diagnostic methods for frontier large language models and multimodal AI systems. You will analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and AI agents, including capability gaps, reasoning errors, robustness issues, and alignment problems. You'll design and build benchmarks and evaluation methods that measure LLM capabilities across text and multimodal modalities, applying post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to data and training interventions that address them. Key responsibilities include conducting root cause analysis on model failures, developing rigorous evaluation frameworks, collaborating with researchers and engineers to define best practices in evaluation-driven AI development, and partnering with top foundation model labs to translate failure analysis into technical and strategic input for next-generation generative AI models. You will publish research findings in top-tier AI conferences. Ideal candidates hold a Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or related field, with deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. You should have hands-on experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development. Published research at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR) and/or journals is expected. Excellent written and verbal communication skills and previous customer-facing experience are valued. Scale works with industry-leading AI labs including Meta, and collaborates with enterprises and government agencies to develop reliable AI systems for critical applications.

Similar roles