SlipstreamJobsFresh Startup & VC-Backed Jobs

Research Scientist – Frontier Evaluations

AfterQuery - San Francisco, CA, United States - In-office - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

AfterQuery is an applied research lab curating data solutions for foundation model development, serving frontier AI labs with the mission of delivering the best data to power the best models. The company is YC's fastest unicorn, valued at $3.2 billion, backed by leading investors including Altos Ventures, BoxGroup, Y Combinator, and angels from Google DeepMind, OpenAI, Anthropic, Meta Superintelligence Labs, and Microsoft AI. As a Research Scientist focused on Frontier Evaluations, you will design and publish rigorous evaluations for frontier AI systems, spanning agentic, coding, and safety evaluations, as well as expert-domain evaluations in applied AI across healthcare, STEM, finance, and related fields. You will own evaluation development end-to-end and collaborate across disciplines to turn important capability gaps into rigorous public research. Key responsibilities include: - Lead the end-to-end design, validation, launch, and continuous improvement of frontier AI benchmarks - Partner with researchers and domain experts to develop evaluations around meaningful model failures, gaps in existing coverage, and high-priority domains - Analyze model capabilities and failure modes using rigorous experimental design and statistical methods - Build reproducible evaluation systems, including harnesses, graders, and benchmark infrastructure - Collaborate with researchers to post-train models and measure the resulting performance gains - Communicate results through benchmark reports, technical articles, and research papers The founding team brings experience from Citadel Securities, Meta, Google, Silver Lake, and Morgan Stanley. You will have the opportunity to shape the engineering organization and lead major technical initiatives as the company scales. REQUIREMENTS: - Strong record of publishing benchmarks or research papers - Clear technical communication and strong scientific writing skills - Commitment to experimental rigor, including baselines, ablations, statistical validity, and contamination controls - Ability to take an ambiguous evaluation question from initial scoping through a reproducible public release - Depth in agentic, coding, and safety evaluations or applied machine learning in an expert domain PREFERRED: - PhD in a related technical field - Research publications at leading conferences or peer-reviewed journals - Interest in multidisciplinary research and the creativity to combine methods and insights from AI, engineering, science, and other expert domains

Similar roles