SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Deepgram is the leading platform for Voice AI, providing real-time APIs for speech-to-text, text-to-speech, and voice agents at scale. The company has processed over 50,000 years of audio and transcribed more than 1 trillion words, serving 200,000+ developers and 1,300+ organizations including Twilio, Cloudflare, and Vapi.
You'll join the team responsible for validating the quality of Deepgram's speech, audio, and multilingual models before they reach customers. This team owns evaluation and quality assurance surfaces ensuring models meet performance targets in both batch and streaming environments. You'll build the pipelines, harnesses, canaries, and test frameworks that catch regressions, hallucinations, and quality issues before they impact customers, partnering closely with Research to turn model expectations into automated, reproducible checks.
In this role, you'll define evaluation methodology and build infrastructure that measures model quality at scale. You'll create evaluation pipelines, define pass/fail criteria grounded in Research benchmarks, and build monitoring that keeps models honest in production. Your work provides trusted signals informing release and optimization decisions, directly protecting customer experience.
Key responsibilities include: designing and building automated evaluation pipelines across batch and streaming (WER, hallucination detection, latency); building scalable evaluation infrastructure with harnesses and orchestration; translating Research benchmarks into automated pass/fail gates; building canaries and continuous-monitoring systems detecting quality regressions in production; partnering with DevOps/Infra on ephemeral test environments; integrating evaluation gates into CI/CD; and raising the bar through code reviews and technical design discussions.
You should have 5+ years of professional software or QA engineering experience with a track record shipping test infrastructure or evaluation systems. Required: BS/MS/PhD in Computer Science, AI, Applied Math, or equivalent; solid backend/scripting experience in Python, Rust, Go, or similar; experience designing automated test pipelines and evaluation frameworks; strong analytical skills reasoning about metrics and statistical variation; ability to tackle ambiguous technical challenges and communicate across research, engineering, and product teams. Nice-to-have: hands-on experience evaluating modern AI systems (LLMs, RAG, agents, multimodal models); React Native or cross-platform mobile framework experience; experience building evaluation frameworks or ML infrastructure for other teams.