SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ServiceTitan is seeking a Director of Software Engineering (Infrastructure) to lead a global Site Reliability Engineering (SRE) team and own the operational excellence of their mission-critical platform serving tens of thousands of trades businesses across North America.
Reporting to the VP of Infrastructure, you will lead, grow, and develop a global SRE team capable of providing 24/7 coverage. Your core responsibilities include operating the operations center and incident command/response functions, achieving and maintaining 4-9's availability across the fleet, and owning release management across the core application and microservices architecture.
Key areas of focus include service capacity planning and demand forecasting, software performance analysis, system tuning, and partnering with development teams to ensure applications are production-ready, scalable, reliable, and observable from day one. You will measure and optimize system performance, identify and implement practices ensuring uptime and reliability, and provide thought leadership on technology matters both internally and externally.
You will establish relationships with peer engineering executives, drive operational best practice adoption across critical services, and represent ServiceTitan in the technology community while assuring customers of continued commitment to their success.
The ideal candidate brings 10-15 years of software engineering experience with a minimum of 7 years leading teams of 50+ engineers. Required expertise includes 7+ years supporting infrastructure in AWS/GCP/Azure, 5+ years delivering enterprise applications in the cloud, 3+ years developing CI/CD pipelines, 3+ years implementing telemetry and observability, and 3+ years establishing SRE practices.
Technical requirements include comprehensive knowledge of cloud services, Infrastructure as Code (Terraform, Ansible), containerization (Docker, Kubernetes), and CI/CD tools (Jenkins). Experience with observability platforms like New Relic, DataDog, and Splunk is essential. A BA/BS in Computer Science or related field is required; advanced degrees are preferred.
Beyond technical skills, you must be a talent magnet with high emotional intelligence, capable of building diverse, inclusive teams where members feel respected and can bring their whole selves to work.