SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Deepgram is the leading platform for Voice AI, providing real-time APIs for speech-to-text, text-to-speech, and voice agents at scale. The company has processed over 50,000 years of audio and serves 200,000+ developers and 1,300+ organizations including Twilio, Cloudflare, and Vapi.
You will own the Text-to-Speech research program end-to-end, setting research strategy, making technical bets, and shipping models to production. This is a hands-on leadership role where you stay deeply involved in the details while building and scaling the team.
Key responsibilities:
- Own the TTS research and model roadmap, deciding which technical directions can materially advance speech-generation quality and recognizing when approaches should pivot or be abandoned.
- Drive advances across neural audio modeling, prosody and expressiveness, controllability, multilingual speech, voice identity and consistency, data and training strategy, post-training, and inference performance.
- Stay deeply technical: review research, challenge assumptions, design experiments, diagnose model failures, and tackle the highest-leverage problems yourself.
- Build evaluation and benchmarking systems that explain why models improve, combining automated metrics with human perceptual assessment.
- Lead a mix of individual contributors and tech lead managers. Hire and develop talent, maintain a high technical bar, grow senior researchers into technical leaders, and set direction across sub-teams while pushing decisions to those closest to the work.
- Partner with engineering and product leadership on production readiness and represent Deepgram's TTS research internally and externally.
The role requires an AI-first mindset. Deepgram operates at the pace of AI with rapid change; you'll be expected to actively use and experiment with advanced AI tools, integrate them into your workflows, and continuously push boundaries. This is not a traditional 9-to-5 role.
REQUIREMENTS (Must-Have):
- Deep expertise in modern TTS, speech generation, or audio generative modeling, with a track record of personally training and improving large-scale neural models.
- Command of the modern speech-generation stack and open problems in naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost.
- Proven history of setting research direction under genuine uncertainty: prioritizing experiments, allocating compute and researcher time, and killing approaches that aren't working.
- Experience leading researchers and research engineers through other technical leaders (developing tech lead managers or equivalent), setting direction across sub-teams, while remaining technically influential.
- AI as your default mode of work, not an occasional tool. You've rebuilt how you and your team operate around AI and have a specific, earned view of its limitations in speech research.
- Ability to make complex technical tradeoffs legible to product, engineering, and executive audiences.
NICE-TO-HAVE:
- TTS or generative-audio models deployed at meaningful production scale.
- Built or substantially scaled a high-performing AI research organization.
- Sophisticated evaluation systems for generative speech; expressive or multilingual generation, voice cloning and adaptation, or controllable generation.
- Recognized external contributions (publications, open source, patents, invited talks) in speech synthesis, neural audio codecs, speech language models, or multimodal models.
- Experience in fast-moving startup or research environments that routinely move models from idea to production.