SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ServiceTitan is seeking a Staff Site Reliability Engineer to own the architecture, design, deployment, and lifecycle management of SQL Server and PostgreSQL databases across Azure, AWS, and on-premises environments. This is a platform-level role with company-wide technical scope, requiring deep expertise in database reliability, performance optimization, and infrastructure automation.
Key responsibilities include:
- Architecting and managing high-availability and disaster recovery solutions (Always On, replication, failover clusters) with defined RTO/RPO objectives
- Leading database design reviews, schema governance, indexing strategies, and query optimization for mission-critical systems
- Implementing comprehensive backup, validation, and recovery procedures
- Performing proactive performance tuning, capacity planning, and workload optimization
- Managing database security including access controls, encryption, auditing, and compliance
- Developing automation for DBA operations using PowerShell, Bash, Python, and Infrastructure as Code
- Establishing monitoring standards and implementing observability using Datadog, Grafana, ELK, and Prometheus
- Participating in incident response, root cause analysis, and postmortem improvements
- Contributing to CI/CD processes for database deployments and schema changes
- Collaborating with engineering teams to optimize data models and resolve production bottlenecks
Required qualifications:
- 8+ years of database engineering/administration experience with demonstrated platform-level strategy ownership
- Deep expertise in performance tuning, high availability/disaster recovery, backup/restore strategies, and database security
- Strong experience with Azure and AWS, including managed services and self-managed deployments
- Proven track record with database migrations (on-prem to cloud, version upgrades, cross-platform)
- Strong scripting skills in PowerShell, Bash, and Python
- Experience implementing monitoring and alerting for database systems
- Solid understanding of reliability engineering principles (SLIs/SLOs)
- Strong troubleshooting skills and ability to manage high-impact production systems
Preferred qualifications include Infrastructure as Code experience, containerized database environments (Kubernetes, Docker), and familiarity with CI/CD pipelines (GitHub Actions, Azure DevOps, TeamCity).