SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Gradle - Remote - Remote - posted 2026-07-31

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Gradle is building Develocity, a toolchain observability and intelligence platform used by Netflix, Airbnb, Spotify, SAP, and hundreds of leading software organizations. The company is AI-native, embedding AI at the core of its product and operations, not as a bolted-on feature. You'll be a founding member of a new SRE team responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services. This is a hands-on role operating production infrastructure at scale. Key responsibilities include: - Operating and maintaining all Develocity instances and supporting services - Participating in follow-the-sun on-call rotation for incident response and troubleshooting across the full stack - Driving automation for deployment, upgrades, monitoring, self-healing, and recovery - Building and maintaining comprehensive observability (logging, metrics, tracing, alerting) - Collaborating with engineering teams to embed reliability into features from the start - Running incident response and retrospectives to drive continuous learning - Owning disaster recovery, backups, and business continuity - Communicating with customers during incidents and maintenance windows - Optimizing performance, resource usage, and costs - Evolving SaaS operations as the company scales You'll work on an internally-built Cloud Application Platform and Kubernetes on AWS, developing deep expertise in both. The team is distributed and remote-first, valuing asynchronous communication and written documentation. Minimum qualifications: 5+ years in SRE, DevOps, or equivalent production operations; strong Kubernetes experience in production; AWS expertise (EKS, RDS, S3, EC2); proficiency with Prometheus, Grafana, and Terraform; incident management track record; SRE best practices knowledge (SLAs, SLOs); scripting proficiency (Python, Bash); 24/7 on-call experience; strong English communication. Preferred: SaaS platform operations at scale, Develocity familiarity, JVM language experience (Java, Kotlin), disaster recovery planning expertise.

Similar roles