SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
commercetools is seeking a Principal Site Reliability Engineer to champion resiliency and operational excellence across the organization. In this role, you will own the discipline of resiliency end-to-end, driving how the company prepares for, responds to, and learns from operational incidents at scale.
Key responsibilities include:
• Standardize Incident Management: Build and champion intuitive, end-to-end processes for incident detection, response, communication, and postmortems across the company.
• Enhance System Visibility: Develop clear, real-time metrics, dashboards, and signals to track system health and incident trends.
• Drive Data-Backed Improvements: Use operational data to identify process gaps and partner with product engineering teams to resolve them.
• Own Peak-Event Readiness: Scale and lead the organization-wide readiness program for massive traffic spikes, such as Black Friday, ensuring mission-critical commerce infrastructure remains stable during high-stakes moments.
• Lead Cross-Team Initiatives: Identify resilience gaps, collaborate with infrastructure and product teams on solutions, and translate ideas into concrete technical outcomes.
• Cross-Functional Collaboration: Partner closely with engineering leadership, Staff Engineers, and domain-specific Principal Engineers across Cloud, Security, API, Architecture, and Performance.
• Foster Knowledge Sharing: Drive organizational communication, documentation, and training around resiliency and operational excellence.
You are a creative problem-solver wired to find solutions in complex challenges and make them simple for others. Your curiosity drives continuous growth and contribution to an environment of trust and teamwork. This is a hybrid role requiring three days per week in one of commercetools' offices: Berlin, London, Munich, or Valencia.