SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Heidi is an AI Care Partner platform supporting clinicians with documentation and care delivery. Founded by clinicians and backed by nearly $100M in funding, Heidi has returned 18+ million hours to clinicians and supported 73+ million patient visits across 116 countries in 18 months. The company partners with major health systems including the NHS, Beth Israel Lahey Health, MaineGeneral, and Monash Health.
You'll join the core Platform/SRE team responsible for production reliability and incident response. This is an intentionally ops-heavy role focused on keeping real systems healthy in production. The team owns incident response, on-call rotations, system reliability, and day-to-day operations for Heidi's platform serving millions of weekly patient visits.
Key responsibilities include:
- Participating in on-call and incident response, responding to production incidents and supporting service restoration with increasing end-to-end ownership over time
- Identifying recurring reliability issues and driving fixes through better alerting, automation, system changes, or process improvements
- Operating and improving Kubernetes clusters, cloud infrastructure, and core platform services with growing ownership
- Strengthening observability through improved dashboards, alerts, logs, and traces focused on actionable signals
- Automating repetitive tasks and simplifying runbooks to reduce operational toil
- Improving deployments, rollback mechanisms, and operational readiness to reduce incident risk from changes
- Writing and maintaining runbooks, participating in blameless post-mortems, and improving incident response processes
- Collaborating with product and feature teams to improve production readiness and reliability expectations
The team operates with a blameless incident culture, prioritizes practical improvements over perfection, and values clear thinking and communication under pressure. Heidi's engineering philosophy emphasizes measured, safe releases guided by data, with stability earning trust and speed delivering impact.
Requirements:
- 3–6+ years in SRE, DevOps, Platform, or operations-heavy engineering roles
- Experience supporting production systems and participating in on-call rotations
- Comfortable debugging live systems under pressure
- Experience operating cloud infrastructure (AWS preferred)
- Working knowledge of Kubernetes and containerized workloads
- Infrastructure as Code experience (Terraform or similar)
- Familiarity with monitoring and alerting tools (Datadog, Prometheus, etc.)
- Scripting or automation experience (Python, Bash, or similar)
Nice to have:
- Experience leading incidents or mentoring others during on-call
- Experience in regulated or security-sensitive environments
- Familiarity with databases, queues, and caches in production
- Interest in reliability practices such as SLOs, error budgets, and capacity planning