SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Attentive is an AI-powered marketing platform specializing in 1:1 personalization across SMS, RCS, email, and push notifications. The company serves over 8,000 customers across 70+ industries, including major brands like Crate and Barrel and Urban Outfitters, and has been recognized by Deloitte's Fast 500, Forbes' Cloud 100, and LinkedIn's Top Startups.
The Platform Infrastructure team is the backbone of Attentive's operations, handling billions of events daily from over 100 million customers. The Production Engineering Team within this organization focuses on delivering a fast, reliable platform that empowers engineers to ship solutions quickly and safely.
As a Staff Site Reliability Engineer, you will take a strategic role designing and implementing solutions that enhance system reliability and scalability. You will mentor other engineers, influence technical roadmaps, and drive cross-functional initiatives across the organization.
Key responsibilities include:
- Designing and implementing systems that enhance reliability, observability, traceability, and incident management
- Leading strategic initiatives and taking ownership of cross-team collaborations
- Partnering with AI/ML, Data, Platform, and Product teams to develop best-in-class services
- Establishing production standards, processes, and tools for operational excellence
- Advocating for and implementing SLIs, SLOs, and reliability-focused metrics across engineering
- Mentoring team members and fostering technical growth
- Driving continuous improvement through creative problem-solving
Required qualifications:
- 7+ years in Production Engineering, Backend Engineering, SRE, DevOps, or similar roles
- Strong technical background with strategic vision beyond immediate problem-solving
- Proficiency in at least one programming language (Golang, Python, Java, TypeScript)
- Demonstrated success delivering medium to large-scale projects improving platform reliability and scalability
- Deep understanding of production reliability concepts (SLIs, SLOs, incident management)
- Excellent communication skills with ability to influence across technical and non-technical teams
- Preferred: experience in dynamic, reliability-focused production environments