SlipstreamJobsFresh Startup & VC-Backed Jobs

Machine Learning Research Scientist, Evaluations

Scale - San Francisco, CA, United States - Hybrid - posted 2026-08-26

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 180,600 - 225,750 / annual

Scale is seeking a Machine Learning Research Scientist to join the Evaluations pod within the GenAI Research Organization. This role focuses on building rigorous benchmarks and diagnostic methods to identify and characterize failure modes in frontier large language models and multimodal AI systems. You will analyze model behavior to uncover capability gaps, reasoning errors, robustness issues, and alignment problems through root cause analysis. You'll design and build evaluation frameworks and benchmarks that measure LLM performance across text and multimodal modalities. Leveraging deep expertise in post-training techniques (SFT, RLHF, reward modeling), you'll connect observed failures to specific data and training interventions that address them. Key responsibilities include: - Identifying and diagnosing failure modes in frontier LLMs and AI agents - Designing benchmarks and evaluation methods for text and multimodal capabilities - Applying post-training knowledge to translate failure analysis into actionable insights - Publishing research findings at top-tier AI conferences - Collaborating with researchers, engineers, and foundation model labs to define best practices in evaluation-driven AI development Ideal candidates hold a Ph.D. or Master's in Computer Science, Machine Learning, AI, or related field, with deep expertise in deep learning, reinforcement learning, and large-scale model fine-tuning. Published research at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR) is highly valued. Experience with RLHF, preference modeling, instruction tuning, and LLM evaluation is essential. Strong communication skills and prior customer-facing experience are preferred. Scale works with leading AI labs to provide high-quality data and accelerate GenAI research. The company partners with industry leaders including Meta, Ernst & Young, Mayo Clinic, and U.S. government agencies.

Similar roles