SlipstreamJobsFresh Startup & VC-Backed Jobs

Lead Site Reliability Engineer

Heidi Health - Melbourne, VIC, Australia - In-office - posted 2026-08-04

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Heidi Health is building AI-powered clinical tools that reduce administrative burden on healthcare providers. The company has achieved rapid scale—supporting 2.5 million patient sessions weekly across 190+ countries—and is now expanding its platform infrastructure. You will lead Heidi's core Platform/SRE team, which owns production systems and reliability. This is a hands-on management role: you'll manage a small team today while staying deeply involved in incident response, on-call rotations, and day-to-day operations. As Heidi scales, you'll grow the team, set operational standards, and represent SRE in engineering leadership conversations. Key responsibilities include: - Participate in on-call and incident response; lead production incidents end-to-end - Improve operational reliability by identifying recurring issues and driving fixes through automation, alerting, and system changes - Own and operate Kubernetes clusters, cloud infrastructure, and core platform services - Strengthen observability by building dashboards, alerts, logs, and traces that surface issues early - Reduce operational toil through automation and improved tooling - Support safe deployments and rollback mechanisms - Contribute to operational practices: runbooks, blameless post-mortems, incident response standards - Collaborate with product and feature teams on production readiness and service ownership - Manage team hiring, onboarding, on-call structure, and career development - Shape team direction, balance reliability work against capacity, and plan with engineering leadership Required qualifications: - 7+ years in SRE, DevOps, platform, or operations-heavy engineering roles, including formal team leadership experience - Track record of hiring, coaching, and growing engineers - Deep production systems experience with on-call rotations; credibility to jump into incidents yourself - Comfortable debugging live systems under pressure - Strong experience with cloud infrastructure at scale (AWS preferred) - Hands-on Kubernetes and containerized workloads in production - Infrastructure as code (Terraform or similar) - Monitoring and alerting tools (Datadog or Prometheus); ability to design alerting strategy - Scripting/automation in Python, Bash, or similar - Experience defining and owning SLOs, error budgets, and capacity planning Nice to have: experience scaling teams through fast headcount growth, regulated/security-sensitive environments, production databases/queues/caches.