SlipstreamJobsFresh Startup & VC-Backed Jobs

Scientific Evals

Edison Scientific - San Francisco, CA, USA - In-office - posted 2026-09-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 170,000 - 220,000 / annual

Edison Scientific builds and deploys AI scientist agents to accelerate science and the development of new medicines. The company is run by scientists and engineers from leading institutions across biology, physics, chemistry, and AI. You will join a team focused on developing rigorous evaluation frameworks, training data, and approaches for advancing AI capabilities in biology. This role sits at the intersection of scientific research, data, and machine learning, ideal for someone with deep scientific training excited to shape how frontier AI systems accelerate discovery. Well-designed benchmarks are critical to measuring abilities and charting progress for AI development. Edison Scientific has been defining the frontier of AI benchmarking for biology research since the first AI scientist agents were introduced, and is seeking people to help push that frontier further. Key responsibilities include: - Design and build novel frontier benchmarks for measuring AI agent-driven science - Develop entirely new evaluation approaches that capture the realities of scientific research, translating the discovery process into evaluable agent tasks - Analyze model or agent work, identify failure modes, and directly contribute to iterative improvements to agents, evaluation tasks, and grading approaches - Collaborate with AI/ML researchers and engineers to translate scientific taste into training signal and agent improvements - Identify and rigorously curate important biological data for improving scientific agents - Coordinate data collection operations and manage workflows, including working with domain experts, tracking progress, and maintaining documentation REQUIREMENTS: - Graduate-level training in biology, computational biology, or related field, with extensive hands-on research experience. Particularly interested in candidates with both wet and dry lab experience. - Understanding or working familiarity with benchmarking and evaluating AI systems - Energized by and willing to own ambitious, open-ended projects requiring creativity, collaboration, and first-principles thinking - Comfortable with Python and able to build workflows for data processing, analysis, and experimentation - Thorough experience working with AI models and tools, including coding agents or Edison's own agent platform, especially in scientific research context - Strong scientific taste/intuition and ability to articulate this into actionable, evaluable workflows - Organized and communicative, able to manage multiple workstreams and coordinate across teams - Detail-oriented and willing to take on high-value but occasionally tedious work PREFERRED QUALIFICATIONS: - Prior experience creating evaluation datasets, annotation guidelines, or working on human-in-the-loop data pipelines - Hands-on experience fine-tuning or evaluating large language models, or familiarity with RLHF and preference-based training - Publications or research experience in areas relevant to AI for science

Similar roles