SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
commercetools is seeking a Principal Site Reliability Engineer to champion resiliency and operational excellence across the organization. In this role, you'll own the discipline of resiliency end-to-end, ensuring the company's mission-critical commerce infrastructure remains reliable at scale—especially during high-stakes moments like Black Friday.
Key responsibilities include:
• Standardize incident management processes across the company, building intuitive workflows for detection, response, communication, and postmortems.
• Enhance system visibility by developing clear, real-time metrics, dashboards, and signals to track system health and incident trends.
• Drive data-backed improvements by analyzing operational data to identify process gaps and partnering with product engineering teams to implement solutions.
• Own peak-event readiness, scaling and leading organization-wide programs to prepare for massive traffic spikes.
• Lead cross-team initiatives to identify resilience gaps and collaborate with infrastructure and product teams on technical solutions.
• Partner closely with engineering leadership, Staff Engineers, and other Principal Engineers (Cloud, Security, API, Architecture, Performance) to drive company-wide resilience improvements.
• Foster knowledge sharing through documentation, training, and organizational communication around operational excellence.
You're a creative problem-solver who thrives in complex environments and excels at making sophisticated challenges simple for others. Your curiosity drives continuous growth, and you bring a collaborative mindset to building trust-based teams. This is a hybrid role requiring three days per week in one of commercetools' offices: Berlin, London, Munich, or Valencia.