SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer

MaintainX - San Francisco, CA, United States - In-office - posted 2026-09-08

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

MaintainX, now part of Autodesk Operations Solutions, is seeking a Site Reliability Engineer to advance platform reliability, observability, and developer autonomy as the company scales. MaintainX is a mobile-first work execution platform serving 13,000+ customers including Duracell, McDonald's, Shell, DHL, and Volvo, managing 13.9 million assets and 79.5 million completed work orders. In this role, you will partner closely with product and platform development teams to improve service stability, resilience, and operational readiness. You'll work across teams to design for reliability from inception, establish clear ownership and standards, and build shared tooling that enables teams to operate services with confidence. You'll contribute to company-wide initiatives defining MaintainX's approach to reliability, including observability standards, incident response practices, and service health metrics, helping the organization adopt proven industry practices at scale. Key responsibilities include assessing service maturity and providing insights to development teams, partnering with teams to implement observability best practices, enabling development teams to become autonomous with service deployment and infrastructure, mentoring developers on reliability practices to foster self-sufficiency, and acting as a bridge between Platform Division teams to drive tooling and practice adoption across the organization. You bring 3–5+ years in software development, SRE, DevOps, or production development roles with hands-on experience operating production systems. You have deep understanding of observability practices in distributed system environments and practical experience with SRE concepts (SLOs, error budgets, incident management). You're proficient in cloud-native platforms and infrastructure-as-code, with working knowledge of at least one programming language (TypeScript/Node.js preferred). Strong communication and collaboration abilities across technical and non-technical teams are essential, along with the ability to translate complex reliability concepts into actionable guidance. You thrive on enabling teams to succeed independently and measure success by reduced dependency on you.

Similar roles