SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Rubrik's Site Reliability Engineering (SRE) team is seeking a Staff Software Engineer - Reliability to serve as a primary technical leader and architect for the company's distributed cloud systems. This role bridges infrastructure strategy, hyperscale automation, and AI infrastructure operations across both global SaaS platforms and government-compliant environments.
Key responsibilities include: formulating architectural vision for Rubrik's Cloud Platform (Kubernetes, MySQL, cloud-native services); building sophisticated internal tools and automation frameworks in Go or Python to eliminate operational toil; deploying and scaling AI infrastructure for multi-tenant SaaS environments; driving AI-driven solutions for SRE and engineering productivity; establishing reliability guardrails as AI adoption accelerates; wielding engineering-wide influence to embed structural resilience across teams; defining and enforcing SLIs, SLOs, and error budgets; serving as primary Incident Commander for high-severity outages; architecting cost-observability and capacity modeling tools; leading the Application-SRE team (a US-based group that partners with engineering, Sales, and Support); and mentoring senior and junior contributors across the organization.
The Application-SRE leadership component is significant: you will set technical direction, track commitments, and position the team as a high-leverage bridge between field signals and engineering roadmaps. You'll participate in on-call rotations and drive durable systemic fixes through blameless post-mortems.
This is a Staff-level individual contributor role with significant organizational influence and people-leadership responsibilities over the Application-SRE team. You'll operate at the intersection of software development and systems engineering, prioritizing automation, self-healing architectures, and structural resiliency.