SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Cantina Labs is a social AI company building advanced real-time generative models for creative expression and social interaction. We're seeking a Research/ML Engineer to join our Speech Team and own the audio side of multimodal generation—from representations through production inference.
You'll architect and ship state-of-the-art speech and audio generation systems with a focus on joint audio-video modeling. This includes designing audio VAEs and neural codecs, building diffusion and flow-matching transformers for large-scale generation, and implementing the conditioning and alignment machinery that keeps characters speaking, singing, and emoting in sync with video. You'll also work on voice cloning, multi-speaker conditioning, cinematic dialogue with music and sound design, and adjacent speech tasks like controllable TTS and voice conversion.
Key responsibilities:
- Design and improve audio representations (VAEs, neural codecs, vocoders) with focus on latent design, reconstruction, and perceptual objectives
- Architect and implement diffusion/flow-matching transformers for audio and video generation, including pre-training, fine-tuning, and alignment (GRPO/DPO)
- Design joint audio-video conditioning and cross-modal alignment mechanisms for multi-shot, multi-speaker generation
- Define data requirements and collaborate on acquisition, curation, AV-sync, quality filtering, and synthetic data strategies
- Design rigorous automated and subjective evaluations for audio fidelity, intelligibility, AV-sync, and robustness
- Drive inference efficiency through distillation, quantization, and kernel optimization to meet latency and cost targets
- Partner with infrastructure on distributed training/inference at scale and production deployment
- Lead small research projects independently while collaborating on larger team initiatives
- Contribute to safety guardrails, watermarking, and misuse mitigation for voice and likeness technology
You see research and engineering as complementary, thrive working across modalities, and are results-oriented and flexible. You enjoy designing experiments, collaborating closely with video, data, and infra teams, and solving unique large-scale problems.