SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff, North Modelling (Evals)

Cohere - London, United Kingdom - Hybrid - posted 2026-08-19

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Cohere is building enterprise AI infrastructure and products, including North, an AI workspace platform for secure, customizable enterprise AI workflows. This role sits at the intersection of applied machine learning, product engineering, and research—focused on building evaluation systems that measure whether AI models are genuinely improving for real customer workflows. You will own the evaluation strategy for North, defining what needs to be measured across agent workflows, tool use, enterprise knowledge work (research, document creation, editing), and human-AI interactions. The core challenge: how do you know if a model is actually getting better for the workflows customers care about? Key responsibilities include: - Designing and owning North's eval strategy across agent workflows and enterprise use cases - Building high-quality evals grounded in real product data: user feedback, production failures, privacy-preserving usage logs, dogfooding, and customer needs - Creating systems that continuously convert user and customer learnings into evals, keeping measurement pace with product evolution - Serving as the voice of North inside the central modeling teams—translating eval results, customer data, and product context into actionable recommendations for model improvements This is a rigorous applied machine learning role for someone who cares deeply about measurement, model behavior, and real-world product quality. You should be excited by the craft of building careful evals: ones that capture messy agentic workflows, reflect actual customer needs, resist benchmark gaming, and provide genuine signal for product and model direction. Ideal candidates have improved LLM-powered or AI-product systems through evals, feedback loops, data curation, or model selection. You care about evaluation as a craft—representative tasks, precise rubrics, clean data, failure analysis, and knowing when a metric is misleading. You have strong applied MLE judgment, can reason clearly about model behavior and production tradeoffs, and are comfortable translating messy qualitative signals into measurement that other teams can act on. You are self-directed, practical, and motivated by open-ended problems requiring both technical depth and product understanding. Cohere is remote-friendly with offices in Toronto, London, NYC, San Francisco, Montreal, Paris, Berlin, and Seoul. This team collaborates across Europe and East Coast North America time zones. Full-time employees receive comprehensive benefits including health/dental, parental leave top-up, 6 weeks paid vacation, education stipends, and home office setup budget.

About Cohere

AI / Data / Infrastructure — enterprise generative AI models and tooling for businesses.

Similar roles