SlipstreamJobsFresh Startup & VC-Backed Jobs

Multimodal AI Model Optimization Research Engineer

Tavus - San Francisco, CA, United States - Hybrid - posted 2026-09-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Tavus is building the human layer of AI, pioneering research in multimodal AI for modeling human-to-human communication (language, audio, video) and generating audio-visual avatar behavior. The company powers text-to-video AI avatars and real-time conversational video experiences across healthcare, recruiting, sales, and education. Series B backed by Sequoia, Y Combinator, and Scale VC. You will join the core AI team as a Research Scientist/Engineer focused on model optimization. Your mission is to take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization techniques. You will own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality. You'll partner closely with researchers and engineers to turn new ideas into deployable systems. The role is ideal for someone who thrives in startup environments, can prioritize independently, and is willing to take calculated risks. You'll work on optimizing diffusion models, video/audio generative models, and large language models in a fast-moving environment. REQUIREMENTS: - Strong experience in deep learning using PyTorch - Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision - Understanding of efficient architectures such as low-rank adapters - Strong understanding of inference performance and GPU/accelerator fundamentals - Strong Python coding skills and reliable research engineering practices - Experience working with large models and datasets in cloud environments - Ability to read ML papers, reproduce results, and adapt ideas - Clear communication and collaboration skills PREFERRED EXPERIENCE: - Optimization of diffusion models, video/audio generative models, or large language models - Experience with real-time or streaming systems (low-latency APIs, WebRTC, streaming TTS/video) - Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA - Experience writing custom Triton/CUDA kernels or low-level performance tuning - Experience with experiment tracking, benchmarking, and profiling at scale - Prior experience in research engineering or applied science roles

Similar roles