SlipstreamJobsFresh Startup & VC-Backed Jobs

Principal Site Reliability Engineer (m/f/x)

commercetools - Berlin, Berlin, Germany - Hybrid - posted 2026-07-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

commercetools is seeking a Principal Site Reliability Engineer to champion resiliency and operational excellence across the organization. In this role, you'll own the discipline of resiliency end-to-end, driving how the company prepares for, responds to, and learns from operational incidents at scale. Your customers rely on commercetools for mission-critical commerce infrastructure, including during peak moments like Black Friday. Key responsibilities include: • Standardize incident management processes across the company—from detection and response through communication and postmortems. • Enhance system visibility by developing real-time metrics, dashboards, and signals to track system health and incident trends. • Drive data-backed improvements by analyzing operational data to identify process gaps and partner with engineering teams on solutions. • Own peak-event readiness, scaling and leading organization-wide programs for massive traffic spikes. • Lead cross-team resilience initiatives, collaborating with infrastructure and product teams to turn ideas into concrete technical outcomes. • Partner closely with engineering leadership, Staff Engineers, and domain-specific Principal Engineers across Cloud, Security, API, Architecture, and Performance. • Foster knowledge sharing through organizational communication, documentation, and training around resiliency and operational excellence. You bring 7+ years of experience driving incident management and operational excellence, with 5+ years leading organization-wide resiliency and reliability initiatives. You have a proven track record managing and scaling systems for high-stakes, multi-team operational events. You're data-literate, able to analyze metrics to diagnose technical issues and measure process improvements. You combine technical depth with leadership and influence, understanding both technical bugs and the organizational dynamics behind them. This is a hybrid role requiring three days per week in the Berlin, London, Munich, or Valencia office.

Similar roles