SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Join Deliveroo's GenAI Platform team within Machine Learning Platform to build shared infrastructure enabling DoorDash, Wolt, and Deliveroo teams to safely deploy GenAI-powered products, agents, automation, and personalization at scale. The team's mission is to increase velocity of business impact from GenAI by running frontier open-weight LLMs and VLMs (GLM, Qwen, Kimi, DeepSeek) with real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs—delivering significant cost and latency wins (e.g., billion embeddings 20× cheaper, visual models 72% cheaper). You'll own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
In this role, you will lead design and architecture of the open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You'll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability while mentoring engineers. You'll architect scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning powering real customer and internal automation use cases. Push the cost and latency frontier of GPU inference—turning multi-day batch jobs into hours and cutting inference costs by multiples—while providing product teams clean choices across open-weight and closed-source models with reliability, fallback, observability, and cost controls. Build platforms supporting rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence. Partner closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to turn emerging GenAI capabilities into durable platform primitives. Set technical direction for the company's centralized GenAI platform including emerging directions such as reinforcement learning (RLHF/RLVR), agent optimization, and post-training and agentic techniques.