SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer - Incident Insights & Readiness

Datadog - Boston, MA, United States - Hybrid - posted 2026-10-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 192,000 - 240,000 / annual

Datadog is seeking a Senior Software Engineer to join the Incident Insights & Readiness SRE team. This team builds software, tooling, and operational frameworks that help Datadog engineers prepare for, respond to, and learn from incidents. You will work across the company to foster a resilient culture by turning incidents into learning opportunities and catalysts for growth. Key responsibilities include: - Own and improve the on-call experience by establishing best practices and building platforms to support on-call rotations and compensation - Define incident response processes and lead the design and implementation of software to streamline incident handling across Datadog - Contribute to the post-mortem process, collaborating on writing and identifying opportunities to reduce friction and enhance organizational learning - Support teams in facilitating incident reviews that emphasize learning and blamelessness, helping share learnings across the organization - Provide technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews and collaborative problem-solving - Train on-callers in incident and post-mortem processes, sharing expertise in incident management best practices - Lead cross-functional initiatives in engineering organizations, embedding with teams to understand challenges and drive improvements to reliability and operational excellence Datadog operates at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. The engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The role is hybrid, allowing you to balance office collaboration with flexible work arrangements. Requirements: - At least 5 years of experience building software that solves real user problems, with experience designing new features and collaborating on code and technical design reviews. Primary development languages are Go and Python, with some TypeScript. - Experience building or operating distributed systems, with familiarity with Kubernetes and understanding of complex failure modes - Demonstrated ability to independently own ambiguous technical problems from design through delivery while balancing long-term engineering quality with pragmatic execution - Experience analyzing incidents, identifying systemic risks, and driving engineering improvements informed by operational learnings - Experience participating in on-call rotations and improving incident response processes. Experience as an incident commander or incident coordinator is a plus. - Strong empathy, collaboration, and communication skills in English to cultivate relationships across various teams - Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority - Welcome candidates from software engineering, site reliability engineering, production engineering, infrastructure, and other roles focused on building reliable systems or improving incident response

Similar roles