SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
LiveKit is building infrastructure for voice-driven AI agents, powering applications for OpenAI, xAI, Salesforce, Spotify, and thousands of others. The company facilitates billions of calls annually and is seeking an exceptional Research Engineer to lead post-training efforts.
In this role, you will own the end-to-end machine learning pipeline for agent training and deployment. Key responsibilities include building training environments and verifiers, managing the synthetic data pipeline from generation through quality gates, running training experiments and analyzing what drives model improvements, designing evaluation frameworks for production releases, selecting and adapting open-weight base models, and ensuring trained behavior works reliably across voice and text channels.
You will ship models into production, monitor their performance on real usage, and iterate continuously. The role requires deep ownership of data quality—treating data as a product with attention to coverage, diversity, and preventing data leakage. You'll design reward systems that resist exploitation and make principled decisions about when to train versus when to use existing models.
You should be a strong Python engineer with a track record of taking models from raw data through production deployment. Experience with post-training techniques (fine-tuning, reward design, reinforcement learning like GRPO) is highly valued. Familiarity with frameworks like TRL, verl, or OpenRLHF, fast inference with vLLM or SGLang, multi-GPU training with FSDP, and open-weight model families (Qwen, Llama) is a plus. Experience building tool-using or multi-turn agents, execution sandboxes, and eval harnesses is also desirable.
You'll work collaboratively with a small, senior team that values craft and creativity in a fully remote environment. The company offers competitive salary and equity, health/dental/vision benefits, and flexible vacation.