SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 140,000 - 180,000 / annual
UJET is an AI-powered contact center platform seeking a Senior Site Reliability Engineer to build and scale a high-impact SRE function. You will be a technical leader responsible for improving system reliability, reducing operational toil, and establishing best practices across the engineering organization.
In this role, you will design how reliability works at UJET, influence engineering decisions, and build the tooling and processes that make production safer and more predictable. You'll lead efforts to improve system reliability, scalability, and performance across critical services. Key responsibilities include defining and implementing SLIs/SLOs and error budgets to guide engineering priorities; designing and developing observability systems (metrics, logging, tracing, alerting) that produce actionable alerts; leading complex incident response as incident commander when needed; conducting postmortems focused on systemic causes and ensuring corrective actions are completed; identifying and eliminating toil through automation and improved workflows; partnering with product and platform teams on architecture decisions and production readiness; building reusable systems and "paved roads" that make it easier for teams to operate services reliably; and mentoring other engineers to raise organizational operational maturity.
Success will be measured by critical services having clear, meaningful SLOs that drive engineering decisions; actionable alerts with reduced noise and manageable on-call workload; efficient incident handling with declining repeat issues; engineering teams adopting reliability best practices with minimal friction; and active toil reduction through automation and better system design.
**Requirements:**
- 6–10+ years of experience in SRE, infrastructure, or backend systems engineering
- Demonstrated experience owning reliability outcomes for complex, distributed systems
- Strong experience with cloud infrastructure (AWS, GCP, or Azure) and production-scale systems
- Deep understanding of observability, incident management, and system performance
- Proficiency in at least one programming language (e.g., Go, Python, Java) with focus on automation and tooling
- Ability to change how other teams work without having managerial authority over them
- Strong competency in making clear decisions during incidents by following a defined process without reacting emotionally
**Preferred Qualifications:**
- Experience building or scaling SRE practices (SLOs, incident frameworks, on-call models)
- Kubernetes/container orchestration experience
- Infrastructure as Code (Terraform, etc.)
- Experience with high-growth or scaling systems
- Background in performance engineering or capacity planning