SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer - Incident Insights & Readiness

Datadog - Paris, France - Hybrid - posted 2026-07-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

The Incident Insights & Readiness SRE team at Datadog builds software, tooling, and operational frameworks that help engineers prepare for, respond to, and learn from incidents. As a Senior Software Engineer on this team, you will own and improve the on-call experience by establishing best practices and building platforms to support on-call rotations and compensation. You'll define incident response processes, lead the design and implementation of software to streamline incident handling, and collaborate with product teams to improve incident response across the organization. You will contribute to the post-mortem process, working with teams on writing and analyzing post-mortems while identifying opportunities to reduce friction and enhance organizational learning. You'll support teams in facilitating incident reviews that emphasize learning and blamelessness, helping share learnings across the organization to improve resilience. A key part of this role involves providing technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews, collaborative problem-solving, and operational excellence best practices. You'll train on-callers in incident and post-mortem processes, sharing expertise in incident management best practices with both newcomers and experienced engineers. You'll lead cross-functional initiatives across engineering organizations, embedding with teams to understand their challenges and drive lasting improvements to reliability and operational excellence. You bring at least 5 years of experience building software that solves real user problems, with expertise in designing features and collaborating on code and technical design reviews. You have experience building or operating distributed systems, familiarity with Kubernetes, and understanding of complex failure modes. You can independently own ambiguous technical problems from design through delivery, balancing long-term engineering quality with pragmatic execution. Experience analyzing incidents, identifying systemic risks, and driving improvements informed by operational learnings is essential. You've participated in on-call rotations and improved incident response processes, with incident commander or coordinator experience being a plus. You demonstrate strong empathy, collaboration, and communication skills in English, with experience mentoring engineers and driving cross-functional initiatives. Datadog welcomes candidates from diverse backgrounds including software engineering, SRE, production engineering, infrastructure, and other roles focused on building reliable systems.

Similar roles