SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Garner Health - Remote - Remote - posted 2026-09-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Garner Health is transforming the U.S. healthcare system by partnering with employers to redesign how healthcare works. Using 550+ proprietary clinical metrics across 80+ specialties and a dataset of 320M+ patients, Garner identifies the best-performing doctors and steers members to higher-quality care, resulting in better outcomes and lower costs. The company has helped 2.5 million people access better care and saved $1B in healthcare costs in five years. As a Senior Site Reliability Engineer on the Platform Engineering team, you will own the reliability, performance, and resilience of Garner's cloud infrastructure powering products and AI/ML workloads. This is an automation-first role where you will define and uphold SLOs, lead incident response, and drive the automation and standards that enable every engineer to ship faster and more reliably. Key responsibilities include: running end-to-end reliability and performance of AWS and Kubernetes cloud environments; serving in on-call rotation and leading incident response with deep-dive root cause analysis; building and maintaining monitoring, alerting, and observability systems; translating scaling requirements into infrastructure-as-code (Terraform) deliverables; proactively identifying cost-efficiency and performance gains; automating operational toil using AI tools; enabling the broader engineering team with deployment and observability standards; and ensuring infrastructure meets security and HIPAA compliance obligations. You will need 4+ years of hands-on production cloud infrastructure experience in SRE, DevOps, or platform engineering; deep expertise with Kubernetes and Terraform in AWS; strong production observability experience including SLO definition and incident response; strong software engineering fundamentals in Python or Go for infrastructure automation; experience with cloud cost-efficiency and performance optimization; and ideally experience supporting AI/ML or data-intensive workloads. Experience in security-conscious or regulated environments (HIPAA, SOC 2) and fluency with AI tools like Claude applied to engineering workflows are valued. The role is based remotely with occasional travel to NYC headquarters.

Similar roles