SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ElevenLabs is seeking a Research Engineer to join its research team, focused on deploying and optimizing frontier AI models in production. The role centers on owning the systems that transform research breakthroughs into real-time products serving millions of users.
Key responsibilities include:
- Deploying state-of-the-art models to production and managing the path from research checkpoint to serving infrastructure
- Optimizing inference performance across the stack (latency, throughput, cost) using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels
- Building and tuning high-performance serving systems for real-time, streaming workloads where millisecond-level performance matters
- Creating tooling and infrastructure that enables researchers to ship new models to production quickly, safely, and with confidence in performance characteristics
The ideal candidate brings:
- Production experience deploying and serving ML models, ideally for latency-sensitive or real-time applications
- Strong GPU programming and inference optimization skills (CUDA, Triton, TensorRT, vLLM, SGLang, or similar frameworks)
- Ability to autonomously profile, diagnose, and eliminate bottlenecks across the serving stack—from model architecture to kernels to orchestration—and build tooling to measure performance
ElevenLabs operates with a high-velocity, impact-driven culture emphasizing rapid experimentation, lean autonomous teams, and minimal bureaucracy. The company uses AI across all functions and prioritizes talent over location. The role is fully remote with optional access to offices in London, New York, San Francisco, and Warsaw. ElevenLabs has raised $781M in funding with a valuation of $11B, backed by Andreessen Horowitz, ICONIQ Growth, and Sequoia.