SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineering Lead

Graphcore - Austin, TX, United States - In-office - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Graphcore is seeking an experienced Site Reliability Engineering Manager to build and lead a new SRE organization responsible for production operations of a rapidly scaling AI supercomputing platform. This is a rare opportunity to establish the reliability function for a new platform from the ground up, taking the SRE organization from initial formation through production launch, stabilization, and scale. You will be responsible for building the SRE team from inception, defining the operating model, establishing production readiness and incident-management practices, and ensuring reliability and operability are engineered into the platform from the outset. This is not a purely managerial position—during development and early production phases, you will work directly with engineering teams, develop deep platform understanding, and participate in troubleshooting and incident response. Key responsibilities include: **Build the SRE Organization:** Hire and develop team members from initial formation through 24x7x365 production operations. Mentor engineers and team leads, develop successors, and build an organization capable of operating effectively without depending on any single individual. Forecast staffing requirements as the platform grows. **Establish the Production Operating Model:** Define the operating model including staffing and coverage, escalation paths, on-call responsibilities, incident management, production access, and change management. Establish clear operational interfaces with Datacenter Operations, engineering teams, vendors, and other service owners. Develop production readiness standards, runbooks, operational procedures, and incident response practices with emphasis on automation. **Engineer Reliability Into the Platform:** Embed with platform engineering teams during development to ensure reliability, serviceability, observability, and operational requirements are incorporated before production. Lead development of SLOs, operational health indicators, alerting standards, and incident severity definitions. Build a culture where recurring operational problems are engineered out through automation and improved design. **Lead Production Operations:** Lead or participate in major production incidents, particularly during platform development and launch. Establish blameless post-incident review processes. Serve as the senior operational authority for the SRE organization and represent production reliability concerns in engineering and leadership discussions. **Requirements:** - Significant experience leading or building an SRE, Production Engineering, Infrastructure Reliability, or comparable function supporting large-scale, highly available production infrastructure - Experience taking a new or rapidly evolving platform through production readiness, launch, stabilization, and ongoing operation - Experience building and operating sustainable 24x7x365 production support or on-call organizations - Strong understanding of modern Site Reliability Engineering principles (SLOs, incident management, observability, automation, toil reduction, capacity management, production readiness) - Strong incident leadership experience managing high-severity, multi-team production incidents under time pressure - Strong systems engineering background with working understanding of Linux, networking, storage, distributed systems, automation, and production infrastructure - Ability to operate effectively at both leadership and hands-on technical levels - Demonstrated ability to lead, hire, mentor, develop, and retain a growing team of strong technical engineers - Strong judgment regarding when to solve immediate operational problems versus investing in eliminating underlying failure modes - Strong written and verbal communication skills, particularly during incidents and when communicating technical risk to senior leadership - Ability to manage conflicting priorities under pressure **Desired Experience (not required):** - High-performance networking technologies (InfiniBand, RDMA, RoCE, large-scale Ethernet fabrics) - Large-scale parallel or distributed storage systems - Workload schedulers and orchestration platforms (Kubernetes, Slurm, or comparable) - Production observability and service-level monitoring for complex distributed infrastructure - Custom or early-generation hardware/firmware environments where hardware and software are developed concurrently - Operational relationships with datacenter operations and infrastructure/hardware/networking vendors - Rapid capacity expansion, datacenter migration, or transitions between temporary and permanent production environments

Similar roles