SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Twilio is seeking a Staff Site Reliability Engineer to join the Platform Engineering organization. SRE owns production health and resiliency at Twilio, accountable for service availability, performance, and recoverability for customers who build their businesses on the platform.
As a Staff SRE, you will work without day-to-day guidance, applying deep subject-matter expertise and industry-leading practices to improve products, processes, and services. Your impact will span multiple teams rather than a single team. This is a hands-on engineering role where you will write and deploy code that improves service reliability, orchestrate complex changes across systems, lead production incident response, and raise the quality bar for engineers around you.
Key responsibilities include:
- Own the reliability posture of production services in your area (availability, latency, capacity, efficiency, performance, monitoring, and alerting)
- Define, instrument, and operate against SLIs and SLOs; use error budgets to drive engineering priorities
- Identify trends and problem areas threatening stability; provide mitigation paths before customer impact
- Drive down repair items and prevent incident classes rather than resolving them one at a time
- Improve detection, response, and recovery; reduce time to acknowledge, engage, mitigate, and restore
- Design for failure: strengthen failure domains, validate recovery paths, make production changes safer to ship and roll back
- Participate in on-call for supported services; lead response when production is degraded
- Write post-mortems identifying true root causes; drive follow-up work to completion
- Oversee efforts to identify, diagnose, report, and document production problems across all reliability dimensions
- Write, configure, and deploy code that measurably improves service reliability (maintainable, reviewed, documented, tested)
- Orchestrate complex changes across systems and services; document design changes, technical decisions, migration plans, and upgrades
- Lead debugging, troubleshooting, and analysis of service architecture and design
- Use code review to drive up quality of coworkers' code
- Reduce operational overhead required to run infrastructure and services
- Drive projects from conception to completion spanning team concerns
- Coordinate across programs; collaborate to estimate and communicate delivery timelines
- Break projects into milestones and tasks; track progress and communicate updates to stakeholders
- Identify and communicate changes that may impact stability
Twilio is a remote-first company. You may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites, or customer meetings. Occasional travel may be required to participate in project or team in-person meetings.
REQUIREMENTS:
Required:
- 8+ years of related engineering experience, with substantial portion focused on reliability, infrastructure, or platform engineering
- Demonstrated accountability for production systems; have carried a pager for services that mattered and owned the outcome when they failed
- Strong software engineering fundamentals; build and ship production code, not only configure tooling
- Experience defining and operating against SLIs and SLOs; using error budgets to inform engineering priorities
- Depth in production operations: incident command, post-mortem analysis, capacity planning, and observability
- Track record of preventing recurrence; reducing incident classes and operational toil, not just closing tickets
- Experience driving changes spanning multiple teams; communication skills to build alignment without formal authority
- Track record of improving engineers around you through code review, design feedback, and mentorship
- Experience with large-scale distributed systems in a cloud environment
Desired:
- Familiarity with infrastructure-as-code, container orchestration, and GitOps-style delivery
- Experience with multi-region architecture, failure-domain design, or regional expansion work
- Background in chaos engineering, game days, or other proactive resilience validation
About Twilio
SaaS / Enterprise Software — authentication and identity infrastructure for developers.