SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Town is building a personalized AI assistant that learns from users' identity, voice, judgment, relationships, and priorities to work across email, calendar, documents, Slack, and other tools. Founded by Jean-Denis Greze (former CTO of Plaid) and Tony Vincent (former Director of Applied AI Product at Google), the company has raised $73M+ from Andreessen Horowitz, Forerunner Ventures, First Round Capital, and Conviction.
As AI Engineer for Evals & Agent Quality, you will own the foundational 0→1 build of Town's evaluation and quality measurement systems. Your core responsibilities include: building a generalized eval framework that measures assistant quality across all surfaces and multi-step agent trajectories; establishing golden datasets and maintaining labeling loops to validate improvements and catch regressions; developing model routing and online evaluation tooling to optimize model selection; instrumenting every prompt and system change to make performance measurable; and partnering across the engineering team to close the loop from quality signals to fixes.
You'll be responsible for setting the bar for what "best" means at Town. The role requires hands-on expertise in LLM evaluation systems, offline/online quality measurement at scale, and a rigorous approach to measurement. You should have strong opinions on eval frameworks and tooling, understand model routing tradeoffs, and be comfortable shipping fixes alongside dashboards and metrics. This is a greenfield opportunity where the system doesn't exist yet, requiring a senior or staff engineer who can operate independently in ambiguity while maintaining high standards for measurement rigor.