SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Dialpad is an AI platform for customer experience that builds agentic AI systems to resolve customer problems in real time across voice and digital channels. The company's AI agents learn from human agents and improve with every interaction, helping organizations deliver better experiences and increase operational efficiency.
As an AI Evaluation Engineer, you will be an integral part of the AI Evaluation team, owning evaluation coverage for Dialpad's Agentic AI systems alongside the existing evaluation lead. This is an individual contributor role reporting to the manager of the AI Evaluation team, based in the Vancouver office.
Key responsibilities include:
- Design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates
- Build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward
- Co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions
- Create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule
- Develop and maintain QA tooling, notebooks, and pipeline components that make recurring evaluations scalable and reusable across teams
- Investigate bugs, triage issues, and decide whether problems should become engineering escalations, test set additions, or follow-up analysis
- Collaborate with cross-functional teams including product, engineering, and data science to support release-readiness decisions
You will focus on LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis. This role is ideal for someone with strong technical skills in evaluation methodology, machine learning, and data analysis who wants to directly impact AI product quality.