SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Hippocratic AI is seeking a Research Scientist to lead speech recognition research and engineering for its healthcare-focused conversational AI platform. This is a senior individual contributor or team-leading role focused on solving the frontier problem of accurate speech recognition in clinical contexts—a challenge where no off-the-shelf solution exists.
In the first 90 days, you will ship measurable improvements to the production ASR system, including improved accuracy on medical terminology, reduced latency, or expanded robustness to diverse patient populations. You'll validate performance gains on clinically relevant benchmarks and establish the data infrastructure roadmap.
By 12 months, you will have designed and deployed a next-generation ASR architecture purpose-built for healthcare, built large-scale medical speech dataset pipelines that create durable competitive advantage, published research at tier-1 venues, and directly shaped how millions of patients experience conversational AI.
Key responsibilities include: designing and developing data-driven ASR models for streaming and non-streaming conversational speech; researching and implementing state-of-the-art speech recognition architectures tailored to the medical domain; training, evaluating, and optimizing ASR models across accuracy, latency, and resource utilization; building data infrastructure and curation pipelines for large-scale medical speech datasets; collaborating with LLM, product, and clinical teams to integrate speech technologies into the platform; and contributing to research culture through rigorous experimentation, documentation, and publications.
You'll work alongside ML researchers, software engineers, and clinicians from Google, Meta, Microsoft, NVIDIA, and Stanford, as well as health system leaders who keep the work grounded in real clinical needs. The role is based in Menlo Park, CA, expected to be five days a week, with potential flexibility for exceptional candidates if a Bellevue presence develops.
Required qualifications: PhD with 3+ years of ASR experience, or Master's with 5+ years of hands-on ASR experience. Must have experience designing and developing algorithms for accurate and efficient speech recognition for both streaming and non-streaming use cases; training, evaluating, and optimizing ASR models; preprocessing and curating large speech datasets; strong Python and C++ programming skills; and comfort with Linux/Unix command-line environments.
Nice-to-have skills include experience building 0-to-1 ASR solutions, hands-on experience with ESPnet, Kaldi, and PyTorch, CUDA experience, leveraging LLMs for enhanced speech recognition, neural/E2E endpointer modeling, and publications in tier-1 journals in speech recognition or NLP.