SlipstreamJobsFresh Startup & VC-Backed Jobs

Machine Learning Engineer, Speech - Joint Audio-Video Modeling

Cantina - Remote - Remote - posted 2026-07-30

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Cantina Labs is a social AI company building advanced real-time generative models for creative expression and social interaction. We're seeking a Research/ML Engineer to join our Speech Team and own the audio side of multimodal generation—from representations through production inference. You'll architect and ship state-of-the-art speech and audio generation systems with a focus on joint audio-video modeling. This includes designing audio VAEs and neural codecs, building diffusion and flow-matching transformers for large-scale generation, and implementing the conditioning and alignment machinery that keeps characters speaking, singing, and emoting in sync with video. You'll also work on voice cloning, multi-speaker conditioning, cinematic dialogue with music and sound design, and adjacent speech tasks like controllable TTS and voice conversion. Key responsibilities: - Design and improve audio representations (VAEs, neural codecs, vocoders) with focus on latent design, reconstruction, and perceptual objectives - Architect and implement diffusion/flow-matching transformers for audio and video generation, including pre-training, fine-tuning, and alignment (GRPO/DPO) - Design joint audio-video conditioning and cross-modal alignment mechanisms for multi-shot, multi-speaker generation - Define data requirements and collaborate on acquisition, curation, AV-sync, quality filtering, and synthetic data strategies - Design rigorous automated and subjective evaluations for audio fidelity, intelligibility, AV-sync, and robustness - Drive inference efficiency through distillation, quantization, and kernel optimization to meet latency and cost targets - Partner with infrastructure on distributed training/inference at scale and production deployment - Lead small research projects independently while collaborating on larger team initiatives - Contribute to safety guardrails, watermarking, and misuse mitigation for voice and likeness technology You see research and engineering as complementary, thrive working across modalities, and are results-oriented and flexible. You enjoy designing experiments, collaborating closely with video, data, and infra teams, and solving unique large-scale problems.

Similar roles