SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior ML Research Scientist, Jockey Core

Twelve Labs - Seoul, South Korea - In-office - posted 2026-08-04

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Twelve Labs builds multimodal AI models that understand video across sight, sound, and motion, powering production-scale AI workloads in media, entertainment, sports, security, and government. The company has raised over $210M from leading investors including NEA, Amazon, NVIDIA, and Snowflake, and operates globally with offices in San Francisco, Seoul, New York, and London. Jockey is Twelve Labs' unified agentic system that reasons across videos and images, combining a reasoning model with a memory layer to build knowledge stores from large video corpora. Unlike traditional models limited by context windows, Jockey decomposes queries, retrieves, segments, and reasons across thousands of videos and images at scale—handling millions of hours of video to deliver corpus-level understanding. The Cognition Models team owns the models that transform video into structured understanding and reasoning, including Pegasus (video-language model) and Jockey Core (the reasoning LLM). The team focuses on multimodal systems with strong instruction-following and complex hierarchical outputs, spanning training infrastructure, temporal segmentation, large-scale inference, data curation, and evaluation pipelines. In this role, you will lead model-efficiency and post-training research for Jockey Core, ensuring a high-quality reasoning model remains efficient enough for production serving without sacrificing quality. Key responsibilities include: driving model compression and efficiency research (structured pruning, quantization, distillation, recovery fine-tuning); designing rigorous evaluations on real agent reasoning and tool-calling behavior using actual traffic; exploring post-training techniques (SFT/RL) to preserve agentic tool-use; and collaborating with serving engineers to translate efficiency gains into real cost and latency improvements. You should have strong LLM research experience in post-training, model compression, distillation, or efficient inference. A track record of independently driving research from ideation through execution with strong experimental judgment is essential. You must be proficient in Python and PyTorch, with the ability to communicate and collaborate effectively across researchers and engineers. Hands-on experience pruning, quantizing, or distilling large models is preferred.

Similar roles