SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Garner Health is transforming the U.S. healthcare system by partnering with employers to redesign healthcare delivery. The company applies 550+ proprietary clinical metrics across 80+ specialties to a dataset of 320M+ patients to identify top-performing doctors and steer members to higher-quality care. In five years, Garner has helped 2.5 million people access better care and saved $1B in healthcare costs. The company recently closed Series E and has doubled five years running.
As Staff Site Reliability Engineer, you will own the reliability strategy for Garner's cloud infrastructure powering products and AI/ML workloads, sitting on the Platform Engineering team. You are the most senior reliability voice in the organization, setting technical direction for how Garner defines, measures, and upholds production quality. Your responsibilities include architecting the SLO framework, incident response program, and automation standards that every engineering team builds upon.
Key responsibilities:
- Architect and own end-to-end reliability, performance, and resilience of Garner's cloud environments (AWS, Kubernetes), including AI/ML workloads; design SLO frameworks for critical services
- Lead the incident response program: set standards for incident response, participate in on-call rotation, lead complex escalations, drive root cause analysis, and build review culture
- Architect the monitoring, alerting, and observability platform to detect and resolve issues proactively
- Transform high-level scaling and reliability requirements into automated, composable infrastructure-as-code (Terraform) deliverables; identify cost-efficiency and performance gains
- Proactively identify inefficient infrastructure workflows and redirect efforts to maximize engineering ROI; convert repetitive operational work into hands-free, monitored processes using AI tools
- Build deployment and observability standards that empower the broader engineering team; mentor engineers across the organization
- Ensure infrastructure meets Garner's security and HIPAA compliance obligations
Ideal candidate has 7+ years hands-on experience operating production cloud infrastructure at scale in SRE, DevOps, or platform engineering roles. Deep expertise with Kubernetes and Terraform in cloud-first environments (AWS preferred), with a track record architecting reliability for systems at scale. Experience designing an organization's reliability practice (SLO frameworks, observability platforms, incident response programs). Strong Python or Go skills for infrastructure automation. Track record driving cloud cost-efficiency and performance optimization. Mentorship experience and ability to set technical direction as senior reliability voice. Excellent communication skills translating complex reliability concepts to technical and non-technical audiences.
Garner is headquartered in NYC but this role is fully remote with occasional travel to HQ.