SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Runway is building world models and generative AI tools for video and creative production. They are seeking an ML infrastructure engineer to own model evaluation end-to-end across the organization.
In this role, you will design and build systems that generate samples at scale, score them with automated metrics and human annotations, track results across checkpoints and releases, and deliver insights to researchers in minutes rather than days. You'll define what "better" means operationally—how the company measures model quality, confidence levels, and ship/no-ship decision criteria.
You'll be embedded within research teams as a member of ML Platform and will set technical direction for evaluation systems across the company. This is a high-leverage role: model quality is bounded by measurement capability.
Key responsibilities include:
- Own the evaluation platform end-to-end: tooling and systems for generating, annotating, reviewing, and adapting at frontier scale
- Define CLIs, APIs, GUIs, and storage layers to make evaluation seamless, fast, and collaborative
- Work directly with research teams on video, image, audio, agents, and robotics to understand measurement needs and generalize solutions into the platform
- Set standards for model evaluation: reproducibility, metric definitions, reporting formats, and trustworthiness thresholds
- Drive broad adoption of the platform across training, production serving, and research efforts
- Contribute to ML Platform team tools and systems for training and serving frontier models
The technical stack includes Python (PyTorch), cloud-native infrastructure, Kubernetes, and agentic development tools. You'll have access to best-in-class models, GPUs, storage, and cloud services.
REQUIREMENTS:
- 5+ years building ML infrastructure or data platforms in production, with experience in evaluation, experimentation, or benchmarking systems
- Strong Python and PyTorch; hands-on experience running large batch GPU workloads on Kubernetes
- Experience designing data pipelines and storage for large volumes of media or model outputs, with attention to versioning and reproducibility
- Familiarity with experimental statistics: paired comparisons, confidence intervals, multiple-comparison pitfalls, inter-rater agreement
- Ability to build internal tools end-to-end (CLI to browser)
- Ability to lead a broad technical area: gather requirements, set direction, make tradeoffs, drive roadmap independently
- Familiarity with full model development lifecycle: data, training, evaluation, serving
- Self-starter who works well embedded with research teams and moves fast
- Strong systems thinking and pragmatic approach to production reliability
- Humility and open-mindedness
NICE TO HAVE:
- Track record building evaluation suites for generative models
- Hands-on work with LLM- or VLM-as-judge pipelines
- Experience with online experimentation platforms
- Prior work evaluating agents or robotics policies
About Runway
AI / Data / Infrastructure; Media / Creator Economy — generative AI tools for video and creative production.