SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Synthesia, the world's leading AI video platform used by over 90% of the Fortune 100, is seeking an ML Platform Engineer to join its ML Platform team in London. The company recently raised $200 million in Series E funding at a $4 billion valuation, backed by premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, and GV.
The ML Platform team builds and operates the systems that enable researchers and product teams to train, serve, and deploy generative models reliably and efficiently. This includes research infrastructure, production serving systems, internal tooling, and platform interfaces. The team is increasingly focused on making these systems automation-friendly and agent-oriented, enabling workflows to be operated through reliable tooling rather than manual effort.
This is a hands-on individual contributor role with significant ownership. You will design and improve platform systems supporting model training, evaluation, and production serving. You'll build infrastructure and tooling that make ML workloads more reliable, scalable, and cost-efficient. Key responsibilities include developing internal tools and workflows operable by both humans and agents, architecting model deployment and serving across research and product environments, improving GPU and cloud infrastructure scheduling and monitoring, and driving improvements in observability, automation, reliability, and developer experience.
You'll collaborate closely with researchers and product engineers to understand pain points and translate them into robust platform capabilities. You'll contribute to technical direction and make pragmatic architectural tradeoffs as the platform scales.
Ideal candidates have strong experience building or operating production systems with focus on reliability, scalability, and maintainability. A systems mindset is essential—thinking naturally in terms of bottlenecks, failure modes, interfaces, resource usage, and long-term operability. Required skills include hands-on cloud infrastructure and Linux experience, Kubernetes for distributed workloads, strong coding in Python or similar backend languages, and experience building internal platforms or developer tooling. Particularly relevant experience includes operating ML infrastructure or model serving systems, supporting research or data-intensive workloads, working with GPU-based systems, and experience with observability in distributed systems. Bonus experience includes building agentic or LLM-powered internal tools, workflow orchestration systems like Temporal, and performance optimization or resource allocation work.