SlipstreamJobsFresh Startup & VC-Backed Jobs

Director of Research, Text to Speech

Deepgram - Remote - Remote - posted 2026-08-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Deepgram is the leading platform for Voice AI, providing real-time APIs for speech-to-text, text-to-speech, and voice agents at scale. The company has processed over 50,000 years of audio and serves 200,000+ developers and 1,300+ organizations including Twilio, Cloudflare, and Vapi. You will own the Text-to-Speech research program end-to-end, setting research strategy, making technical bets, and shipping production models. This is a hands-on leadership role where you stay deeply technical while building and leading a high-performing research team. Key responsibilities include: owning the TTS research and model roadmap, deciding which technical directions materially advance speech-generation quality, driving advances across neural audio modeling, prosody, expressiveness, controllability, multilingual speech, voice identity, data strategy, and inference performance. You'll build evaluation and benchmarking systems that explain why models improve, not just whether they did. You'll lead a mix of individual contributors and tech lead managers, hiring and developing both while maintaining an exceptionally high technical bar. You must have deep expertise in modern TTS, speech generation, or audio generative modeling with a track record of personally training and improving large-scale neural models. You need command of the modern speech-generation stack and open problems around naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost. You must have experience setting research direction under uncertainty, prioritizing experiments, allocating compute and researcher time, and killing approaches that aren't working. Leadership experience managing researchers and research engineers through technical leaders is essential, as is the ability to make complex technical tradeoffs legible to product, engineering, and executive audiences. AI should be your default mode of work. You've rebuilt how you and your team operate around AI tools and have an earned view of what AI still can't do in speech research. Nice-to-haves include TTS models deployed at production scale, experience building or scaling high-performing AI research organizations, sophisticated evaluation systems for generative speech, recognized external contributions (publications, open source, patents, talks), and experience in fast-moving startup or research environments.

Similar roles