SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Deepgram is the leading platform for Voice AI, providing real-time APIs for speech-to-text, text-to-speech, and voice agents at scale. The company has processed over 50,000 years of audio and transcribed over 1 trillion words, serving 200,000+ developers and 1,300+ organizations including Twilio, Cloudflare, and Vapi.
You'll join as a Senior Data Scientist sitting at the intersection of research and data, working on conversational audio—a domain orders of magnitude more complex than text. Speech carries speakers, accents, dialects, emotion, overlapping talk, code-switching, and acoustic conditions ranging from quiet studios to drive-throughs, all with rich context.
Key responsibilities include: characterizing Deepgram's conversational data (languages, conditions, domains, quality, underrepresentation); designing and building active-learning loops to systematically decide what to work on next; thinking deeply about human-in-the-loop workflows and tooling; building representative benchmarking datasets and methodology; running experiments that separate what actually works from assumptions; ensuring data consistency, cleanliness, and accessibility; and bringing method and automation to model adaptation.
You'll own high-leverage, hands-on work with unusual latitude to define an area from first principles, reporting to the VP of Data Operations. The role rewards conviction, creativity, and willingness to overturn assumptions when data says otherwise.
Required: hands-on experience with real data pipelines and model-facing problems in data science, ML, or applied research; strong Python and data tooling; experience with data characterization, data selection, or active learning; working familiarity with speech/audio or NLP models; track record of turning messy data into measurable model or product improvements; ability to build reusable systems; strong communication skills; active AI-tool user.
Desirable: direct ASR/TTS or audio data experience; multilingual/code-switched data work; ensemble labeling, pseudo-labeling, or LLM-assisted annotation; data provenance, PII/GDPR-aware pipelines, or model-improvement compliance.
Deepgram operates with an AI-first mindset—AI use and comfort are core to how the company operates. Change is rapid, and you'll need to experiment, adapt, think on your feet, and learn constantly.