SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Twelve Labs is building AI infrastructure to enable machines to understand video at scale. The company has raised over $300M from top-tier investors (NEA, Radical Ventures, Amazon, NVIDIA, Snowflake, Databricks, Index Ventures, and others) and operates globally from offices in San Francisco, Seoul, New York, and London. Core R&D happens in Seoul.
The Construction team (Video Ingestion & Serving Platform) develops end-to-end systems that process customer video from upload through to a form usable by AI models for search, analysis, and generation. The team handles video collection, decoding, embedding, metadata processing, storage, and model serving—operating reliably in both large-scale multi-tenant SaaS and enterprise deployment environments. The team uses Python and Go services, Temporal, PostgreSQL, Kubernetes, KServe/vLLM, Terraform, ArgoCD, and Grafana/OpenTelemetry.
You will lead the Construction team and raise engineering execution across the broader Jockey Engineering and Research organization. You own the team's composition, execution, technical direction, and production outcomes. You will build high-performing teams, clarify priorities, support engineer growth, and guide complex platform projects from problem definition through production operations.
This is a technically deep engineering manager role. You will drive major architecture and operational decisions, participate in design and code reviews, and directly contribute to problem-solving and implementation when needed. You must be able to trace production issues to their root cause—whether in services, databases, workflows, Kubernetes, or model-serving infrastructure—and lead the team through resolution.
Beyond Construction, you will collaborate with other engineering managers and leaders in Infrastructure, Backend, ML, and Research to improve CI/CD, release stability, observability, incident response, developer productivity, testing, and AI-driven development practices across Jockey. Early cross-org impact will come through technical trust, collaboration, and voluntary adoption by other teams rather than direct organizational authority.
Key responsibilities include: leading the Construction team through clear goals, feedback, performance management, hiring, and career development; establishing technical direction for large-scale video processing, embedding, storage, and AI model-serving platforms in collaboration with Backend, ML, Infrastructure, Research, and Product teams; improving durable workflows and data systems (concurrency, retries, idempotency, backpressure, failure recovery, PostgreSQL scaling, safe live migration); advancing Kubernetes and infrastructure-as-code scalability with the Infrastructure team; optimizing throughput, latency, GPU utilization, reliability, and cost through load testing, profiling, and operational metrics; raising operational readiness through SLOs, observability, alerting, on-call practices, and blameless incident response; and improving engineering effectiveness across Jockey through better CI/CD, automated testing, release/rollback stability, developer productivity, service ownership, and AI-driven development standards.
QUALIFICATIONS
Required:
- Substantial software engineering experience combined with proven track record building and managing teams responsible for production systems, including hiring, coaching, performance management, and career development
- Experience designing and operating distributed systems (workflows, queues, databases, storage, large-scale data processing) and leading safe live migrations
- Go or Python proficiency at a level where you can lead code reviews and implementation decisions, and contribute directly when needed
- Experience operating Kubernetes and cloud infrastructure as code, and improving performance, reliability, or cost through load testing, profiling, monitoring, and operational data
- Ability to drive multi-team decisions and production rollouts through technical trust and influence, not just reporting authority
- Judgment to define ambiguous problems, communicate clearly, and balance execution speed against reliability, maintainability, and long-term platform health
Preferred:
- Experience building or operating GPU-based inference serving systems (KServe, vLLM, Triton)
- Experience with durable workflow engines (Temporal, Cadence, Step Functions)
- Experience with sharded or distributed databases (Aurora Limitless, Citus, Vitess, Cassandra, FoundationDB)
- Experience building and operating observability stacks (Grafana, Mimir, Loki, Alloy, OpenTelemetry)
- Experience with FFmpeg, video decoding, transcoding, or large-scale media processing pipelines
- Experience improving CI/CD, developer platforms, developer productivity, or operational practices across multiple teams
- Experience adopting and operating AI-driven development practices alongside quality, security, and maintainability standards
- Strong English communication skills for seamless collaboration with global teams