SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer

GYANT - Remote - Remote

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 135,000 - 160,000 / annual

Fabric Health is seeking a Site Reliability Engineer to architect and evolve the AWS and Kubernetes (EKS) infrastructure powering healthcare experiences for millions of patients. You will design, deploy, and maintain enterprise-grade Kubernetes clusters while optimizing core AWS services (EC2, RDS, S3) for performance, cost, and reliability. Key responsibilities include: **Infrastructure & Kubernetes Orchestration**: Design and maintain EKS clusters for high availability. Optimize infrastructure footprint across AWS services. Build scalable infrastructure management platforms with clear interfaces and guardrails. Define golden paths for workload orchestration and integration with infrastructure dependencies. **Automation & AI-Assisted Operations**: Develop robust, reusable GitHub Actions components for the software delivery lifecycle. Create internal tools that replace manual operations with intelligent, autonomous systems. Explore and deploy agentic workflows for AI-assisted runbooks that automate complex procedures and repetitive tasks. **Observability & Incident Management**: Drive observability practices by maintaining reliable mechanisms to collect metrics, traces, and logs. Lead incident response efforts and facilitate blameless postmortems to systematically reduce MTTR. Define and monitor platform-level SLIs and SLOs to meet rigorous healthcare performance standards. **Compliance & Collaboration**: Ensure continuous HIPAA compliance across all infrastructure. Review architectural decisions and technical proposals to surface risks early. Mentor engineers on reliability best practices and contribute clinical-safety perspectives to cross-functional design reviews. You should design systems and shape processes with a focus on empowering colleagues. You dig for root causes, use AI tools judiciously, and thrive in fast-paced, safety-critical environments balancing pragmatism with technical rigor. You prefer building systems others can run independently rather than being the indispensable go-to person. Required: 5+ years SRE or Platform Engineering experience managing production environments at scale. Expert technical depth in AWS (EKS, EC2, RDS, S3) and production-grade Kubernetes. Proficiency with Terraform, Datadog, Helm, and GitHub Actions. Solid coding/scripting skills in Python, Bash, or Go. Rigor-first mindset with dedication to HIPAA-compliant, high-availability architecture. Preferred: Experience building agentic workflows or AI-assisted tooling.

Similar roles