SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Benchling is an AI platform for biotech R&D used by over 200,000 scientists globally, including major biopharma companies like Sanofi and Moderna. The company is building AI agents and models to accelerate scientific discovery and compress decades of R&D work into years.
This role is part of a team focused on making frontier AI models better at science. While large language models possess extensive biological knowledge, significant gaps remain in reasoning for real-world scientific problems. You'll work at the intersection of software engineering, biology, and frontier AI to close this gap.
Key responsibilities include building high-quality datasets for evaluating and improving frontier models by transforming complex scientific data into rigorous tasks and environments for LLMs. You'll analyze model failure modes by running experiments across frontier models to identify where they struggle and opportunities for improvement. You'll design and build scalable data infrastructure with pipelines that curate, transform, and validate large volumes of scientific data into model evaluation and improvement tasks.
You'll collaborate directly with frontier AI labs to develop and evaluate new approaches for improving models on challenging scientific tasks, and work closely with scientists to translate expert judgment into problems and evaluation criteria that reliably distinguish strong model behavior.
This is an early-stage, rapidly evolving area where the right technical approaches are still being discovered. The team works in-person in San Francisco Monday through Friday in a fast-paced environment where priorities can shift and rapid experimentation is encouraged.