SlipstreamJobsFresh Startup & VC-Backed Jobs

Applied Research Scientist, AI Research

Descript - San Francisco, CA, USA - Hybrid - posted 2026-08-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 261,625 - 299,000 / annual

Descript's Research team builds generative video and audio models and multimodal understanding systems that power distinctive features like Video Regenerate, lipsync, video translation, and zero-shot voice and roomtone cloning. This is applied research—everything ships to production and reaches millions of creators within months. You'll work on generative media synthesis, building models for photorealistic video and audio synthesis at production scale. You'll develop zero-shot voice and roomtone cloning from minimal reference audio, create multimodal vision-language systems for agentic editing features, and tackle computer-vision-heavy problems including digital human reconstruction and facial modeling. You'll also design new algorithms for media synthesis, speech enhancement, anomaly detection, and audio/video tagging. A key part of the role is direction-setting: identifying and pursuing the next research direction that should become a Descript feature, not just a paper. More senior candidates will own this directly; more junior candidates will grow into it. You'll need proven ability to design and implement deep learning algorithms, demonstrated through publications, open-source work, or shipped models. Strong programming skills and deep fluency in PyTorch and/or TensorFlow are required. You must have a track record of generating new ideas in machine learning—producing more ideas than you can implement and running many experiments quickly once infrastructure is established. Strong experimental judgment is essential: testing ideas fast and being honest about which don't work. A PhD or Master's in deep learning or related field is expected, or equivalent demonstrated experience. At minimum, you must have either led or been first author on an accepted publication at a top venue (NeurIPS, ICML, ICLR, ICASSP, ICCV, CVPR, Interspeech, SIGGRAPH, or similar) or played a key role in shipping a production feature with deep learning as a core component. The team is small and senior-heavy, running lean by design. You'll own problems end-to-end—framing, experiments, and analysis—rather than handing pieces off. Clear writing and communication are essential, especially when a direction isn't working. Everyone is expected to generate ideas faster than any one person could implement them, then pick well and move efficiently. Depth in generative modeling for video/audio/images, vision-language models, speech/audio modeling, building evaluation systems for generative outputs, or taking research ideas through to shipped production features is a strong signal, though not required.

Similar roles