SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Gradle - Remote - Remote - posted 2026-07-31

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Gradle is building Develocity, a toolchain observability and intelligence platform used by leading software organizations including Netflix, Airbnb, Spotify, and SAP. The company is AI-native, embedding AI at the core of its product and operations rather than as a bolted-on feature. You'll be a founding member of Gradle's new SRE team, responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services. This role operates production infrastructure at scale, including internally-built cloud platforms, Kubernetes on AWS, and supporting services like artifact registries. Key responsibilities include operating and maintaining all Develocity instances, participating in follow-the-sun on-call rotations with incident response ownership, driving automation across deployment, upgrades, monitoring, and recovery. You'll build and maintain comprehensive observability (logging, metrics, tracing, alerting), collaborate with engineering teams to embed reliability into features, run incident response and retrospectives, and own disaster recovery and business continuity planning. You'll also optimize performance and resource usage while communicating with customers during incidents. The team is distributed and remote-first, emphasizing asynchronous communication and written documentation. You'll need strong self-direction and clear cross-timezone communication skills. Minimum qualifications: 5+ years in SRE, DevOps, or equivalent production operations; strong Kubernetes experience in production; AWS expertise (EKS, RDS, S3, EC2); proficiency with Prometheus, Grafana, and Terraform; incident management track record; SRE best practices knowledge (SLAs, SLOs); scripting proficiency (Python, Bash); 24/7 on-call experience; strong English communication. Preferred: SaaS platform operations at scale, Develocity familiarity, JVM language experience (Java, Kotlin), disaster recovery planning expertise.

Similar roles