SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Gradle is building Develocity, a toolchain observability and intelligence platform used by Netflix, Airbnb, Spotify, SAP, and hundreds of leading software organizations. The company is AI-native, embedding AI at the core of its product and operations, not as a bolted-on feature.
You'll be a founding member of a new SRE team responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services. This is a hands-on role operating production infrastructure at scale.
Key responsibilities include:
- Operating and maintaining all Develocity instances and supporting services
- Participating in follow-the-sun on-call rotation for incident response and troubleshooting across the full stack
- Driving automation for deployment, upgrades, monitoring, self-healing, and recovery
- Building and maintaining comprehensive observability (logging, metrics, tracing, alerting)
- Collaborating with engineering teams to embed reliability into features from the start
- Running incident response and retrospectives to drive continuous learning
- Owning disaster recovery, backups, and business continuity
- Communicating with customers during incidents and maintenance windows
- Optimizing performance, resource usage, and costs
- Evolving SaaS operations as the company scales
You'll work on an internally-built Cloud Application Platform and Kubernetes on AWS, developing deep expertise in both. The team is distributed and remote-first, valuing asynchronous communication and written documentation.
Minimum qualifications: 5+ years in SRE, DevOps, or equivalent production operations; strong Kubernetes experience in production; AWS expertise (EKS, RDS, S3, EC2); proficiency with Prometheus, Grafana, and Terraform; incident management track record; SRE best practices knowledge (SLAs, SLOs); scripting proficiency (Python, Bash); 24/7 on-call experience; strong English communication.
Preferred: SaaS platform operations at scale, Develocity familiarity, JVM language experience (Java, Kotlin), disaster recovery planning expertise.