SlipstreamJobsFresh Startup & VC-Backed Jobs

Machine Learning Engineer - Voice Conversion

Cantina - Remote - Remote - posted 2026-08-04

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 200,000 - 220,000 / annual

Cantina, founded by Sean Parker, is building an advanced AI character creator platform that enables lifelike AI bots to interact across voice, video, and text. The company is creating infinitely scalable, personalized content experiences and seamless group chat functionality. You'll join the Speech Team as a Research/ML Engineer to build state-of-the-art speech systems end-to-end. This role bridges cutting-edge research and practical engineering, focusing on voice conversion, controllable text-to-speech (TTS), voice design, and related tasks. You'll own the model-data-evaluation flywheel, partnering closely with research, data, and infrastructure teams to ship fast, reliable, and cost-aware models. Key responsibilities include: architecting and implementing large-scale audio models (>8B parameters) with diffusion/flow-matching transformers; designing and running rigorous experiments with listening tests and objective metrics; developing tooling to enhance team productivity; owning data requirements from acquisition through curation and synthetic data strategies; designing automated evaluations including robustness and bias checks; hardening training-evaluation-inference pipelines for production SLAs; and contributing to safety guardrails and misuse mitigation. You should have exceptional hands-on experience with large-scale audio models, deep expertise in diffusion and flow-matching transformers, strong background with audio VAEs and neural codecs, proficiency in multi-node distributed training (FSDP/DeepSpeed), and proven ability to ship generative models to production. Strong PyTorch and performance optimization skills (CUDA/Triton) are essential. Experience with voice cloning, speech control, or expressive speech generation is highly valued. The ideal candidate sees research and engineering as complementary, is results-oriented and flexible, enjoys designing experiments and metrics, and thrives on solving large-scale problems.

Similar roles