SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 175,000 - 200,000 / annual
KAYAK, part of Booking Holdings, is a leading travel search engine serving billions of queries across global metasearch brands including momondo, Cheapflights, and HotelsCombined. This role leads the team responsible for the reliability, scalability, and security of KAYAK's container orchestration platform at global scale.
You will combine deep technical expertise with people leadership, guiding a team of platform engineers while shaping the strategy for containerized production environments. You'll be a key partner to engineering, security, and infrastructure teams, enabling KAYAK to move faster and more safely.
Key responsibilities include:
- Own the technical roadmap and operational strategy for the Kubernetes platform, ensuring cluster reliability, scalability, and security across all environments
- Lead a team of container platform engineers with full responsibility for hiring, performance management, and career development
- Deliver continuous improvements through cluster lifecycle management (version upgrades, node provisioning, capacity planning, workload migration from legacy platforms)
- Coach engineers on Kubernetes best practices, container security, RBAC design, and operational excellence
- Collaborate with software engineering, security, and infrastructure teams to ensure reliable, observable, deployable containerized workloads
- Drive adoption of infrastructure-as-code and GitOps practices to reduce manual toil and improve consistency
- Design and enforce standards for Kubernetes observability (metrics, alerting, dashboarding) for rapid incident detection and resolution
- Maintain high standards for operational excellence including on-call practices, incident response, and blameless post-mortems
- Align the Kubernetes platform roadmap with broader infrastructure and product engineering goals
- Advance security posture through cluster hardening, network policy enforcement, and compliance partnerships
Benefits include remote flexibility (up to 20 days/year), company-paid mental health support (therapy and Headspace), company-wide annual week off, no-meeting Fridays, paid parental leave, generous PTO, paid volunteer time, development dollars, leadership development, travel discounts, employee resource groups, competitive retirement and health plans, and free lunch twice weekly.
REQUIREMENTS:
- 8+ years in SRE, DevOps, or platform/infrastructure engineering, with at least 3 years in engineering leadership or management
- Proven track record operating large-scale, 24/7 production Kubernetes environments including cluster lifecycle management and workload migration from legacy orchestration
- Deep, hands-on Kubernetes expertise: RBAC design, network policy, cluster hardening, multi-environment operations (on-premises and cloud)
- Strong technical foundation in on-premises and cloud infrastructure (GCP and/or AWS) and infrastructure-as-code tools (Terraform, Ansible)
- Experience as a people manager with demonstrated ability to hire, develop, and retain engineering talent; commitment to inclusive, high-performing teams
- Solid experience with modern observability tooling (Prometheus, Grafana, or equivalent) including alerting, dashboarding, and incident-driven tuning
- Proficiency in at least one scripting/automation language (Python, Go, or Bash)
- Excellent communication skills with ability to articulate clear technical vision and work effectively across engineering, security, and product teams
- Degree in Computer Science or related technical field, or equivalent practical experience