SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Pivotal Health is a technology platform helping healthcare providers navigate complex reimbursement landscapes and recover fair payment from insurance companies. The company combines software, data, and AI-driven tools to simplify Independent Dispute Resolution (IDR) workflows and reduce administrative burden on stretched provider teams.
You'll join as a Staff Site Reliability Engineer—a senior individual contributor role with broad influence across engineering. This is hands-on infrastructure work paired with technical leadership responsibility. You'll define and strengthen reliability, scalability, and operational excellence across Pivotal's platform as it scales from rapid growth into a predictable, resilient system meeting healthcare's high standards for availability, security, and auditability.
Key responsibilities include: setting Pivotal's reliability strategy and roadmap; designing resilient, scalable cloud infrastructure; building world-class observability (metrics, logs, traces, dashboards, alerting); improving incident response and disaster recovery practices; reducing operational toil through automation; embedding reliability practices across engineering teams; strengthening security and compliance for sensitive healthcare and financial data; and providing technical leadership to senior engineers and engineering leaders.
You'll partner closely with software, data, AI, security, and product teams to design systems that recover gracefully, improve observability, and reduce operational risk. You'll establish service-level objectives, guide architecture decisions, mentor engineers, and raise the bar for infrastructure and operational engineering across the organization.
Required: 8+ years in SRE, infrastructure, or platform engineering operating large-scale production systems. Deep expertise in cloud infrastructure, distributed systems, networking, containers, orchestration, IaC, and modern deployment practices. Proven ability designing and operating highly available systems in fast-growing environments. Strong grasp of observability, SLOs, capacity planning, incident management, disaster recovery, and performance engineering. Hands-on debugging across application, infrastructure, network, and data layers. Experience building automation and internal tooling. Comfortable influencing architecture across teams without formal authority. Pragmatic communicator translating operational risk and technical tradeoffs for diverse stakeholders. Energized by greenfield work and operating in ambiguity.
Bonus: experience with healthcare, financial, or regulated data infrastructure.