SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 158,000 - 237,000 / annual
Rubrik is seeking a Site Reliability Engineer to join its Engineering team, focused on ensuring the reliability and performance of Rubrik's infrastructure services, particularly supporting the FedRAMP-certified Polaris Cloud Platform.
In this role, you will be responsible for maintaining high availability and durability of Rubrik's databases, designing and implementing relational database systems for performance and reliability, and managing backend systems including Kubernetes and MySQL. You'll establish best practices for internal teams to write performant SQL queries, perform periodic database upgrades while minimizing customer downtime, and drive reliability, availability, and efficiency improvements across the platform.
Key responsibilities include:
- Ensure high availability and durability of databases
- Design, implement, and maintain relational database systems
- Manage and operate backend systems like Kubernetes and MySQL
- Perform database upgrades with minimal downtime
- Participate in on-call rotations using a follow-the-sun model across continents
- Write and review code, plan and execute upgrades, develop documentation and capacity plans
- Debug production issues and work cross-functionally with various engineering teams
- Build monitoring tools and automation to increase team efficiency
- Drive the FedRAMP certification process
- Support large-scale SaaS infrastructure serving enterprise customers with high SLO and SLA requirements
This position carries special security and privacy responsibilities for protecting U.S. Federal Government interests, including compliance with system-specific security policies, data protection requirements, role-based training, and collaboration with Information Security on security controls.
REQUIREMENTS:
- 2+ years of experience designing and managing relational databases with focus on performance, scalability, reliability, high-availability, and disaster recovery
- Experience in database design and architecture supporting large enterprise customers with high SLO and SLA requirements
- Experience operating the database layer of a large-scale SaaS product
- Proficiency in one or more programming languages: Golang, Python, Java, Scala, or C++
- Systematic problem-solving approach with strong communication skills and sense of ownership
- Expertise in designing, analyzing, and troubleshooting large-scale distributed systems
- Ability to debug and optimize code and automate routine tasks
- Strong operational experience with Unix/Linux operating systems and networking
- Experience with Google Cloud Platform or other public cloud technologies
- Minimum 1-3 years of experience as a Development, DevOps, or Site Reliability Engineer
- Willingness to provide 24/7 coverage
- Strong documentation skills
- Experience working with multiple departments and divisions
- Strong understanding of databases
- Experience leading support personnel
- Experience with FedRAMP certification strongly desired
- U.S. citizenship at time of hire (required for FedRAMP program)
- Residence within the contiguous United States (lower 48 states and District of Columbia)
- Willingness to undergo a Single Source Background Investigation if required