SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
commercetools is seeking a Principal Site Reliability Engineer to champion resiliency and operational excellence across the organization. In this role, you'll own the discipline of resiliency end-to-end: mature incident management, strong operational visibility, data-driven process improvement, and organization-wide readiness for peak-traffic events.
Key responsibilities include:
• Standardize Incident Management: Build and champion intuitive, end-to-end processes for incident detection, response, communication, and postmortems across the company.
• Enhance System Visibility: Develop clear, real-time metrics, dashboards, and signals to track system health and incident trends.
• Drive Data-Backed Improvements: Use operational data to find process gaps, partnering with product engineering teams to fix them.
• Own Peak-Event Readiness: Scale and lead the organization-wide readiness program for massive traffic spikes like Black Friday.
• Lead Cross-Team Initiatives: Identify resilience gaps, collaborate with infrastructure and product teams on solutions, and turn ideas into concrete technical outcomes.
• Cross-Functional Collaboration: Partner closely with engineering leadership, Staff Engineers, and domain-specific Principal Engineers (Cloud, Security, API, Architecture, Performance).
• Foster Knowledge Sharing: Drive organizational communication, documentation, and training around resiliency and operational excellence.
You bring 7+ years driving incident management and operational excellence, with 5+ years leading organization-wide resiliency and reliability initiatives. You have a proven track record managing and scaling systems for high-stakes, multi-team operational events. You're data-literate, able to analyze metrics to diagnose technical issues and measure process improvements. You combine technical depth with leadership and influence, evaluating both technical bugs and the organizational/human dynamics behind them.
This is a hybrid role requiring three days per week in the Berlin, London, Munich, or Valencia office.