SlipstreamJobsFresh Startup & VC-Backed Jobs

Director, Text-to-Speech Synthesis Research

Deepgram - Remote - Remote - posted 2026-09-17

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Deepgram is the leading platform for Voice AI, providing real-time APIs for speech-to-text, text-to-speech, and voice agents at scale. The company has processed over 50,000 years of audio and serves 200,000+ developers and 1,300+ organizations including Twilio, Cloudflare, and Vapi. You will own the Text-to-Speech research program end-to-end, setting research strategy, making technical bets, and shipping models to production. This is a hands-on leadership role where you stay deeply involved in the details while building and scaling the team. Key responsibilities: - Own the TTS research and model roadmap, deciding which technical directions can materially advance speech-generation quality and recognizing when approaches should pivot or be abandoned. - Drive advances across neural audio modeling, prosody and expressiveness, controllability, multilingual speech, voice identity and consistency, data and training strategy, post-training, and inference performance. - Stay deeply technical: review research, challenge assumptions, design experiments, diagnose model failures, and tackle the highest-leverage problems yourself. - Build evaluation and benchmarking systems that explain why models improve, combining automated metrics with human perceptual assessment. - Lead a mix of individual contributors and tech lead managers. Hire and develop talent, maintain a high technical bar, grow senior researchers into technical leaders, and set direction across sub-teams while pushing decisions to those closest to the work. - Partner with engineering and product leadership on production readiness and represent Deepgram's TTS research internally and externally. The role requires an AI-first mindset. Deepgram operates at the pace of AI with rapid change; you'll be expected to actively use and experiment with advanced AI tools, integrate them into your workflows, and continuously push boundaries. This is not a traditional 9-to-5 role. REQUIREMENTS (Must-Have): - Deep expertise in modern TTS, speech generation, or audio generative modeling, with a track record of personally training and improving large-scale neural models. - Command of the modern speech-generation stack and open problems in naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost. - Proven history of setting research direction under genuine uncertainty: prioritizing experiments, allocating compute and researcher time, and killing approaches that aren't working. - Experience leading researchers and research engineers through other technical leaders (developing tech lead managers or equivalent), setting direction across sub-teams, while remaining technically influential. - AI as your default mode of work, not an occasional tool. You've rebuilt how you and your team operate around AI and have a specific, earned view of its limitations in speech research. - Ability to make complex technical tradeoffs legible to product, engineering, and executive audiences. NICE-TO-HAVE: - TTS or generative-audio models deployed at meaningful production scale. - Built or substantially scaled a high-performing AI research organization. - Sophisticated evaluation systems for generative speech; expressive or multilingual generation, voice cloning and adaptation, or controllable generation. - Recognized external contributions (publications, open source, patents, invited talks) in speech synthesis, neural audio codecs, speech language models, or multimodal models. - Experience in fast-moving startup or research environments that routinely move models from idea to production.

Similar roles