SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 150,000 - 300,000 / annual
Turing is seeking exceptional AI Research Scientists to join its STEM research organization and develop new ways to evaluate, train, and improve frontier AI systems. This is a research-first role focused on problems where the right benchmark, dataset, or methodology often does not yet exist. You will identify important gaps in the literature, propose ambitious new research directions, and take projects from initial hypothesis through experimentation, benchmark construction, and publication.
The research team focuses on frontier STEM evaluation, synthetic data, hallucination and reliability, and agentic science. You will work on four primary areas:
**Frontier Benchmarks and Evaluation**: Identify high-impact gaps in existing benchmark and evaluation literature. Design novel benchmarks in and across STEM fields and on general model functionality. Develop evaluations for emerging model capabilities poorly captured by traditional static benchmarks. Design rigorous task-generation, grading, contamination-control, difficulty-calibration, and validation methodologies. Build benchmarks that become both valuable research contributions and meaningful standards for evaluating frontier models.
**Synthetic Data and Post-Training**: Develop methods for generating high-quality synthetic STEM training data. Study how task selection, difficulty, diversity, verification, filtering, and data quality affect downstream performance. Explore methods for generating useful training signal in domains where expert human data is scarce or expensive. Design experiments that determine when synthetic data genuinely improves capabilities.
**Hallucination, Reliability, and Verification**: Study hallucination, uncertainty, calibration, and epistemic failure in technical domains. Develop evaluations and methods for improving factual reliability, self-correction, verification, citation, and appropriate abstention. Investigate when models should reason internally, invoke tools, seek external evidence, or recognize knowledge gaps.
**Agentic Science**: Research AI systems capable of performing extended scientific and technical work. Develop workflows involving literature search, coding, simulation, tool use, experimentation, verification, and iterative reasoning. Evaluate long-horizon scientific agents and identify bottlenecks. Explore new approaches to human-AI and multi-agent scientific collaboration.
You will have significant latitude to propose new research programs in reasoning, model evaluation, AI-for-science, data generation, and emerging capabilities. Success means identifying major capabilities that existing benchmarks fail to measure, discovering failure modes in synthetic-data pipelines, building evaluations that change how frontier labs understand key phenomena, or developing agentic workflows that advance performance on complex scientific tasks.
Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks. The company also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment.
This role is required to be in office five days a week, based in any of Turing's offices in San Francisco, Palo Alto, or Seattle.
**Requirements**:
- PhD or equivalent research experience in machine learning, computer science, mathematics, physics, chemistry, biology, engineering, statistics, or another highly technical field
- Demonstrated ability to formulate and execute original research
- Strong understanding of modern LLMs and the frontier AI research landscape
- Excellent experimental design, quantitative reasoning, and scientific judgment
- Ability to rapidly understand unfamiliar technical literature and develop expertise in new areas
- Strong Python skills and the ability to independently build research prototypes and evaluation pipelines
- Excellent technical writing and communication
- Comfort working in a fast-moving environment where the research agenda evolves with the frontier
- Strong publication record valued, but primary focus is on ability to identify important questions, design rigorous ways to answer them, and execute quickly