SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
KAYAK, part of Booking Holdings, is seeking a Senior MLOps Engineer to design and implement machine learning infrastructure and production lifecycle systems. This is a hands-on senior role bridging data science and production engineering, based in the Berlin office with a requirement to commute 3 days per week.
You will join the Machine Learning Platform team, responsible for building and maintaining scalable infrastructure and automated pipelines for model training, deployment, and monitoring. Your focus will be ensuring ML models are reliable, reproducible, and performant at scale.
Key responsibilities include:
- Build and maintain end-to-end ML infrastructure: Extend and operate infrastructure powering model deployment, including CI/CD pipelines, model orchestration, and automated training pipelines designed to scale reliably without manual intervention.
- Own model deployment and serving: Define and evolve standards and tooling for model serving, ensuring low latency and high availability across ML services.
- Develop core MLOps capabilities: Establish and maintain infrastructure functioning as reliable, self-service systems for the ML organization, including feature stores, model registries, and automated monitoring for performance and data drift.
- Operationalize infrastructure for the ML team: Collaborate with Operations to enable Kubernetes autoscaling and GPU provisioning as accessible self-service tools, standing up and operating a Kubernetes-based development cluster and taking models from experimentation to GPU-backed production.
- Improve platform reliability and performance: Partner with Operations to design resilient monitoring using advanced observability tooling, define service-level objectives, and implement automation to reduce manual interventions.
- Empower Data Scientists: Build standardized, optimized workflows and "golden paths" that streamline the model development lifecycle, letting Data Scientists focus on modeling while you handle infrastructure.
Required qualifications: Experience building and operating ML platforms in production environments; solid knowledge of containerization and orchestration (Docker, Kubernetes), Linux internals, and model serving at scale; familiarity with ML lifecycle tooling including orchestration frameworks, feature stores, model registries, and drift/performance monitoring; experience owning production systems with SLOs, observability tools (Prometheus, Grafana, Datadog), incident response, and systematic failure diagnosis; production-quality code writing in Python or comparable language; experience modernizing production infrastructure with attention to reliability, risk, and cost; ability to take ownership of technical outcomes, advocate using data, and communicate clearly to technical and non-technical audiences.