SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Heidi Health is building AI-powered clinical tools that reduce administrative burden on healthcare providers. The company has achieved rapid scale—supporting 2.5 million patient sessions weekly across 190+ countries—and is now expanding its platform infrastructure.
You will lead Heidi's core Platform/SRE team, which owns production systems and reliability. This is a hands-on management role: you'll manage a small team today while staying deeply involved in incident response, on-call rotations, and day-to-day operations. As Heidi scales, you'll grow the team, set operational standards, and represent SRE in engineering leadership conversations.
Key responsibilities include:
- Participate in on-call and incident response; lead production incidents end-to-end
- Improve operational reliability by identifying recurring issues and driving fixes through automation, alerting, and system changes
- Own and operate Kubernetes clusters, cloud infrastructure, and core platform services
- Strengthen observability by building dashboards, alerts, logs, and traces that surface issues early
- Reduce operational toil through automation and improved tooling
- Support safe deployments and rollback mechanisms
- Contribute to operational practices: runbooks, blameless post-mortems, incident response standards
- Collaborate with product and feature teams on production readiness and service ownership
- Manage team hiring, onboarding, on-call structure, and career development
- Shape team direction, balance reliability work against capacity, and plan with engineering leadership
Required qualifications:
- 7+ years in SRE, DevOps, platform, or operations-heavy engineering roles, including formal team leadership experience
- Track record of hiring, coaching, and growing engineers
- Deep production systems experience with on-call rotations; credibility to jump into incidents yourself
- Comfortable debugging live systems under pressure
- Strong experience with cloud infrastructure at scale (AWS preferred)
- Hands-on Kubernetes and containerized workloads in production
- Infrastructure as code (Terraform or similar)
- Monitoring and alerting tools (Datadog or Prometheus); ability to design alerting strategy
- Scripting/automation in Python, Bash, or similar
- Experience defining and owning SLOs, error budgets, and capacity planning
Nice to have: experience scaling teams through fast headcount growth, regulated/security-sensitive environments, production databases/queues/caches.