SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
KAYAK, part of Booking Holdings, is seeking a Senior MLOps Engineer to design and implement machine learning infrastructure and production lifecycle systems. This is a hands-on senior role bridging data science and production engineering, based in the Berlin office with hybrid flexibility (3 days/week on-site).
You will join the Machine Learning Platform team, responsible for building and maintaining scalable infrastructure and automated pipelines for model training, deployment, and monitoring. Key responsibilities include:
- Build and maintain end-to-end ML infrastructure: extend and operate CI/CD pipelines, model orchestration, and automated training pipelines designed to scale reliably without manual intervention.
- Own model deployment and serving: define and evolve standards and tooling for model serving, ensuring low latency and high availability across ML services.
- Develop core MLOps capabilities: establish and maintain feature stores, model registries, and automated monitoring for performance and data drift.
- Operationalize infrastructure for the ML team: collaborate with Operations to enable Kubernetes autoscaling and GPU provisioning as self-service tools, including standing up and operating a Kubernetes-based development cluster.
- Improve platform reliability and performance: partner with Operations to design resilient monitoring using advanced observability tooling, define service-level objectives, and implement automation to reduce manual interventions.
- Empower Data Scientists: build standardized, optimized workflows and "golden paths" that streamline the model development lifecycle.
Required qualifications: experience building and operating ML platforms in production; solid knowledge of containerization and orchestration (Docker, Kubernetes), Linux internals, and model serving at scale; familiarity with ML lifecycle tooling including orchestration frameworks, feature stores, model registries, and drift/performance monitoring; experience owning production systems with SLOs, observability tools (Prometheus, Grafana, Datadog), incident response, and large-scale failure diagnosis; production-quality code in Python or comparable language; experience modernizing production infrastructure with attention to reliability, risk, and cost; strong ownership, data-driven decision-making, and clear communication skills.