SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Machine Learning Engineer, Jockey Core

Twelve Labs - Seoul, South Korea - In-office - posted 2026-08-04

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Twelve Labs builds multimodal AI models that understand video at production scale across media, entertainment, sports, security, and government. The company has raised $210M+ from top-tier investors including NEA, Amazon, NVIDIA, Snowflake, and Databricks, with offices in San Francisco, Seoul, New York, and London. Jockey is Twelve Labs' unified agentic system that reasons across video and image corpora at scale—handling millions of hours of video. Unlike traditional models limited by context windows, Jockey decomposes queries, retrieves relevant segments, and reasons across thousands of videos to deliver timestamped, actionable results. The system is built for both human users and autonomous AI agents, with production-grade infrastructure designed for reliability and scale. Jockey Core is the reasoning LLM at the heart of Jockey. It sits in the critical path of every agent step, directly shaping quality, latency, and cost. The Cognition Models team owns the models that turn video into structured understanding—including Pegasus (video-language model) and Jockey Core itself. The team spans training infrastructure, temporal segmentation, large-scale inference, data curation, and evaluation pipelines, working with advanced compute including NVIDIA B300s and Blackwell. In this role, you will lead serving engineering for Jockey Core from engine selection through production scale-out. You'll build benchmarks and load tests that replay real agent traffic (short prompts, frequent round-trips, long tool-output contexts), measuring TTFT and inter-token latency. You'll apply inference optimization techniques—quantization, batching/scheduling, disaggregated prefill/decode, speculative decoding—to hit cost and latency targets. You'll build cost models and drive production hardening through autoscaling, capacity planning, observability, and failover mechanisms via rollback-safe rollouts. You'll collaborate to ensure model-efficiency gains translate to real serving wins and set the serving technical bar through design review. The role also includes exploring AI-assisted development tools to improve productivity. You should have significant production experience serving and optimizing large-scale LLM inference (vLLM, TensorRT-LLM, SGLang, or similar), with expertise in batching/scheduling, quantization, disaggregated prefill/decode, and speculative decoding. Experience designing and operating large-scale distributed systems in high-performance GPU environments is essential. You'll drive ambiguous technical decisions with measured latency/throughput/cost data and have a track record building observability, SLOs, and failure-response for production services.

Similar roles