SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Gradle Inc. - Remote - Remote

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 180,000 - 205,000 / annual

Gradle Inc. is building Develocity, a toolchain observability and intelligence platform used by leading software organizations including Netflix, Airbnb, Spotify, and SAP. The company is AI-native, embedding AI at the core of its product and operations rather than as a bolted-on feature. Develocity helps software teams achieve delivery excellence through deep observability, build and test acceleration, and AI-powered intelligence across the entire toolchain, supporting Gradle Build Tool, Apache Maven, sbt, npm, Python, and Bazel. You will be a founding member of a new SRE team, responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services. You'll work on the company's internally-built Cloud Application Platform and Kubernetes on AWS, developing deep expertise in these systems. Your responsibilities include operating and maintaining all Develocity instances and supporting services; participating in an on-call rotation to own incident response and troubleshooting across the stack; driving automation across application deployment, upgrades, monitoring, self-healing, and recovery; building and maintaining observability for all managed services (logging, metrics, tracing, and alerting); collaborating with engineering teams to build reliability into features from the start; running incident response and retrospectives to drive continuous learning; owning disaster recovery, backups, and business continuity; communicating with customers during incidents and maintenance windows; optimizing performance, resource usage, and costs; and helping evolve SaaS operations as the company grows. The team is distributed and remote-first, valuing asynchronous communication and written documentation. Strong self-direction and clear communication across time zones are essential. You'll have real ownership of production systems used by engineers at major companies, direct interaction with customers, and the opportunity to shape how SRE practices are established in a growing organization. REQUIREMENTS: Minimum Qualifications: - 5+ years in SRE, DevOps, or equivalent role operating production services at scale - Strong Kubernetes experience in production environments - Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2) - Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform) - Track record of incident management and response - Knowledge of SRE best practices (SLAs, SLOs) - Scripting proficiency (Python, Bash) for automation - Experience with 24/7 on-call rotations - Strong written and verbal English communication Preferred Qualifications: - Experience operating SaaS platforms at scale - Familiarity with Develocity - JVM language experience (Java, Kotlin) - Disaster recovery planning and execution experience - Customer-facing incident communication skills - Experience establishing SRE practices in new or growing teams

Similar roles