SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Gradle is building Develocity, a toolchain observability and intelligence platform used by Netflix, Airbnb, Spotify, SAP, and hundreds of leading software organizations. The company is AI-native, embedding AI across its product to help teams achieve delivery excellence through deep observability, build acceleration, and AI-powered intelligence.
You will be a founding member of a new SRE team, serving as a technical and operational leader for reliability across Develocity's production services. This is a hands-on, influential role that shapes how the company operates—defining SRE vision, setting operational standards, and mentoring team members as the function scales.
Responsibilities include operating and maintaining all Develocity instances and supporting services; defining and evolving SRE standards, practices, and operating models (on-call, incident response, postmortems, SLOs); participating in follow-the-sun on-call rotations as a technical escalation point; leading incident response and blameless retrospectives; setting reliability priorities using risk, customer impact, and error budgets; identifying systemic reliability risks; leading architectural and design reviews; driving automation across deployment, monitoring, and operational workflows; building comprehensive observability (logging, metrics, tracing, alerting); owning disaster recovery and business continuity; partnering with engineering leadership to balance feature delivery with reliability; mentoring and coaching SREs; and contributing to hiring and onboarding.
You'll work on an internally-built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in the full stack—from application to infrastructure. The team is distributed, remote-first, and values asynchronous communication and written documentation. Strong self-direction and clear cross-timezone communication are essential.