SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Machine Learning Ops Engineer

KAYAK - Berlin, Berlin, Germany - Hybrid - posted 2026-08-07

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

KAYAK, part of Booking Holdings, is seeking a Senior Machine Learning Ops Engineer to design and implement machine learning infrastructure and production lifecycle systems. This is a hands-on senior role bridging data science and production engineering, based in the Berlin office with a requirement to commute 3 days per week. You will join the Machine Learning Platform team, responsible for building and maintaining scalable infrastructure and automated pipelines for model training, deployment, and monitoring. Your focus will be ensuring ML models are reliable, reproducible, and performant at scale. Key responsibilities include: - Building and maintaining end-to-end ML infrastructure: CI/CD pipelines, model orchestration, and automated training pipelines designed to scale reliably without manual intervention - Owning model deployment and serving standards and tooling to ensure low latency and high availability across ML services - Developing core MLOps capabilities including feature stores, model registries, and automated monitoring for performance and data drift - Collaborating with Operations to enable Kubernetes autoscaling and GPU provisioning as self-service tools, including standing up and operating a Kubernetes-based development cluster - Improving platform reliability and performance through resilient monitoring, service-level objectives, and automation to reduce manual interventions - Empowering Data Scientists by building standardized, optimized workflows and clear "golden paths" for the model development lifecycle Required qualifications: - Production experience building and operating ML platforms - Solid knowledge of containerization and orchestration (Docker, Kubernetes), Linux internals, and model serving at scale - Familiarity with ML lifecycle tooling: orchestration frameworks, feature stores, model registries, drift/performance monitoring - Experience owning production systems: defining SLOs, building observability (Prometheus, Grafana, Datadog), incident response, and diagnosing large-scale failures - Production-quality code writing in Python or comparable language - Experience modernizing production infrastructure with attention to reliability, risk, and cost - Strong ownership mentality, data-driven decision-making, and clear communication with technical and non-technical audiences Benefits include 6 weeks paid vacation, paid parental leave, company-paid mental health support (therapy and Headspace), annual company-wide week off, no meeting Fridays, development dollars, leadership development, travel discounts, and office perks including free lunch twice weekly and bike leasing.

Similar roles