SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Engineering Manager (Resilience)

Monday.com - Warsaw, Poland - In-office - posted 2026-08-26

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: PLN 73,000 - 80,000 / monthly

Monday.com is seeking a Senior Engineering Manager to lead the Resilience Group, a critical team responsible for ensuring platform availability, incident response speed, and reducing production incidents across their AI-powered work platform serving 250,000+ customers with over $1B in ARR. The Resilience Group operates across three core areas: Traffic Shaping (tooling for handling unusual usage patterns and traffic spikes), Observability (enabling engineers to detect and troubleshoot issues), and Reliability Engineering (providing practices and tools for service operations). As the manager of this group, you will own the mission and outcomes of platform availability and incident response, working closely with a Group Tech Lead to define technical strategy and translate it into actionable work scopes. Key responsibilities include: - Owning platform availability, incident response speed, and reducing production accidents - Identifying bottlenecks across the platform and leading systematic solutions (e.g., improving observability coverage and detection rates) - Managing and developing Team Leaders, supporting their growth and ensuring cross-team alignment - Defining and improving incident response practices, including tooling, escalation paths, post-mortems, and preventative mechanisms - Understanding platform architecture, usage patterns, and product context to identify high-impact problems - Working with the Tech Lead to translate technical strategy into executable work The Warsaw office, established in 2022, is a growing engineering hub working on impactful infrastructure and platform challenges. The team works on problems ranging from detecting traffic anomalies at scale to managing database servers and content security. REQUIREMENTS: - Strong distributed systems understanding: ability to navigate microservices environments, understand infrastructure and application-level reliability, and assess trade-offs across observability, safety mechanisms, and operational tooling - Deep background in SRE or production engineering with excitement about emerging practices (agentic SRE, agentic incident response, AI-assisted operations) - Senior engineering management experience: proven track record managing Team Leaders and Senior Technical Leaders in complex technical domains - Problem-solving mindset with understanding of system architecture, organizational constraints, and ability to implement engineering-creative solutions that ship - Familiarity with observability best practices, traffic management, and production reliability at scale - Experience owning end-to-end reliability programs (not just running ceremonies, but defining what needs to be built and why)

Similar roles