SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff, Evals

Handshake - San Francisco, CA, United States - Hybrid - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Handshake AI partners with frontier AI labs to organize expert human knowledge and advance the AI economy. The company builds systems that turn expert knowledge into data and evaluations that improve frontier models. As a Member of Technical Staff, Evals, you will define how frontier AI systems are measured, understood, and improved. This is a high-ownership role for researchers who build. You will partner with AI researchers, domain experts, and customers to develop new benchmarks, reward and verifier systems, agent-evaluation methodologies, and data-quality techniques. Key responsibilities: - Design and build evaluation frameworks, benchmarks, and methodologies for frontier LLMs, AI agents, multimodal models, and reinforcement-learning environments - Develop reward models, programmatic verifiers, graders, and other feedback systems that make model behavior measurable and improvable - Research what makes evaluations representative, difficult, reliable, and resistant to shortcutting or reward hacking - Build systems for high-quality human data, including expert task design, annotation methodologies, data-quality signals, and data-attribution techniques - Run fast, rigorous iteration loops: prototype, evaluate, interpret results, diagnose failure modes, and turn learnings into the next benchmark or system - Publicly contribute to the field through benchmarks, open-source tools, research, and technical writing You will work alongside engineers, researchers, operators, and builders from Scale AI, Meta, Google, Amazon, xAI, Notion, and Palantir. Early members of the team will have unusual influence over technical direction, operating culture, and the open-source software, benchmarks, and research products built. REQUIREMENTS: - PhD in ML/AI, computer science, data science, or related fields (or equivalent research experience in industry) - Publications at top AI/ML venues such as NeurIPS, ICML, ICLR, COLM - Demonstrated builder mindset with experience tinkering with agents and shipping high-quality software, benchmarks, or datasets (e.g., strong GitHub profile, OSS contributions, or product portfolio) - Strong Python skills and experience building scalable software and working with agents - Strong knowledge of frontier AI: benchmarks, eval techniques, agent harnesses, post-training recipes, data shapes - Comfort operating in an ambiguous, fast-moving environment with substantial ownership

Similar roles