SlipstreamJobsFresh Startup & VC-Backed Jobs

Principal Site Reliability Engineer (m/f/x)

commercetools - Munich, Bavaria, Germany - Hybrid - posted 2026-08-20

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

commercetools is seeking a Principal Site Reliability Engineer to champion resiliency and operational excellence across the organization. In this role, you'll own the discipline of resiliency end-to-end: mature incident management, strong operational visibility, data-driven process improvement, and organization-wide readiness for peak-traffic events. Key responsibilities include: • Standardize Incident Management: Build and champion intuitive, end-to-end processes for incident detection, response, communication, and postmortems across the company. • Enhance System Visibility: Develop clear, real-time metrics, dashboards, and signals to track system health and incident trends. • Drive Data-Backed Improvements: Use operational data to find process gaps, partnering with product engineering teams to fix them. • Own Peak-Event Readiness: Scale and lead the organization-wide readiness program for massive traffic spikes like Black Friday. • Lead Cross-Team Initiatives: Identify resilience gaps, collaborate with infrastructure and product teams on solutions, and turn ideas into concrete technical outcomes. • Cross-Functional Collaboration: Partner closely with engineering leadership, Staff Engineers, and domain-specific Principal Engineers (Cloud, Security, API, Architecture, Performance). • Foster Knowledge Sharing: Drive organizational communication, documentation, and training around resiliency and operational excellence. You bring 7+ years driving incident management and operational excellence, with 5+ years leading organization-wide resiliency and reliability initiatives. You have a proven track record managing and scaling systems for high-stakes, multi-team operational events. You're data-literate, able to analyze metrics to diagnose technical issues and measure process improvements. You combine technical depth with leadership and influence, evaluating both technical bugs and the organizational/human dynamics behind them. This is a hybrid role requiring three days per week in the Berlin, London, Munich, or Valencia office.

Similar roles