SlipstreamJobsFresh Startup & VC-Backed Jobs

Engineering Manager, Edge SRE

Cloudflare - London, United Kingdom - Hybrid - posted 2026-10-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Cloudflare is seeking an Engineering Manager to lead the Edge SRE team based in London. You will manage and develop a team of Site Reliability Engineers responsible for maintaining Cloudflare's edge production infrastructure and building tools that enable all engineering teams to understand, debug, and optimize production systems. The Edge SRE team is part of Cloudflare's Infrastructure Engineering organization, which operates globally across Asia, Europe, and the US to provide follow-the-sun production reliability coverage. SREs work closely with all engineering teams at Cloudflare, who participate in on-call schedules for their services. The SRE organization prioritizes incident remediation and follow-up work above product innovation, and production incident impact directly influences team priorities. In this role, you will lead engineers focused on edge distribution where most client traffic is served. You will be responsible for mentoring and growing your team, helping them develop skills and confidence to make independent decisions. You'll work with individuals on your team to build personal development plans aligned with Cloudflare's goals, take an active role in prioritizing the SRE organization's roadmap, and drive cross-team and cross-organizational alignment across engineering, infrastructure, and product teams. Key responsibilities include leading the team to keep Cloudflare's edge reliable and scalable, mentoring and empowering engineers, participating in deep technical design discussions, partnering with other Engineering Managers to achieve reliability outcomes, and ensuring high-quality systems are built. You will also participate in incident root cause analysis and follow-ups, and help mature the tooling and best practices that enable engineering teams to debug in production, measure availability and performance indicators, and track thresholds. Desirable experience includes hands-on software or reliability engineering background, experience leading and hiring teams that build and run tools and platforms, strong planning and execution skills with predictability, incident management expertise, comfort managing teams with deadlines and short release cycles, familiarity with observability tools (Jaeger, OpenTracing, ELK, Prometheus, Thanos, Grafana, Clickhouse), experience running and maturing distributed systems, familiarity with proxies, DNS, databases, internet protocols and security, and experience developing tools and APIs. REQUIREMENTS: - 5+ years of software engineering, reliability, or operations experience in a customer-focused environment - 2+ years of experience managing a team of 5 or more engineers on projects in distributed systems, tooling, Linux, internetworking, infrastructure security, or infrastructure management - Comfortable collaborating and coordinating on cross-team projects and workflows - Ability to provide strong technical vision for systems and infrastructure teams - Experience building services and systems, taking projects from inception to production, and providing leadership for major projects - Capable of leading discussions with upper management and tailoring technical detail to suit audience

Similar roles