SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Hippocratic AI is seeking a Staff Site Reliability Engineer to own the design and operation of a GPU management and scheduling platform that sits at the center of their healthcare AI infrastructure. The company runs nearly 30 models across heterogeneous hardware, and this role focuses on the complex engineering challenge of keeping that fleet fast, reliable, and cost-effective.
You will design and build the GPU management and scheduling platform that decides when, where, and how inference calls run across the fleet. Key responsibilities include: building the metrics pipeline that collects GPU load and utilization data and translates those signals into real-time decisions; implementing admission control to protect capacity by deciding when to accept, queue, or shed inference requests; building autoscaling systems that adjust model replicas in response to demand; developing cloud orchestration systems and operators in Python and Go to manage the fleet; architecting and operating scalable, fault-tolerant, secure production systems on AWS, GCP, or Azure; designing infrastructure automation and deployment pipelines (Terraform, CI/CD) as first-class software; standing up monitoring, logging, and alerting systems; developing and enforcing security and compliance policies for a healthcare AI platform; partnering with engineers and research scientists to diagnose and resolve complex infrastructure issues; and mentoring engineers to raise the technical bar across the team.
Hippocratic AI is building the world's first healthcare-only, safety-focused LLM platform. The company was co-founded by CEO Munjal Shah and a team of physicians, hospital leaders, and AI pioneers from institutions including El Camino Health, Johns Hopkins, Washington University in St. Louis, Stanford, Google, Meta, Microsoft, and NVIDIA. The company recently raised a $126M Series C at a $3.5B valuation, bringing total funding to $404M with participation from leading healthcare and AI investors including CapitalG, General Catalyst, a16z, Kleiner Perkins, and others.
REQUIREMENTS
Must-Have:
- 10+ years of professional experience across site reliability / DevOps engineering and software engineering
- Computer Science degree required from a top CS program
- Strong software engineering fundamentals — building orchestration and scheduling systems in Python and/or Go, not just configuring off-the-shelf tools
- Experience designing systems that make decisions from operational metrics — collecting signals, interpreting them, and driving control loops such as autoscaling, load shedding, or admission control
- Deep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD, or similar)
- Hands-on production experience with at least one major cloud platform (AWS, GCP, or Azure)
- Strong knowledge of containerization and orchestration (Docker, Kubernetes)
- Experience with monitoring and logging stacks (ELK, Grafana, Datadog, or similar)
- Familiarity with secrets management and security tooling (HashiCorp Vault, AWS KMS, Azure Key Vault)
- Excellent problem-solving skills and ability to work both independently and collaboratively
- Strong communication and interpersonal skills
Nice-to-Have:
- Experience managing GPU fleets or scheduling workloads across heterogeneous accelerators
- Familiarity with ML inference serving and model deployment (e.g., Triton, KServe, Ray Serve, or similar)
- Experience with Kubernetes autoscaling internals (HPA/VPA, custom metrics, custom controllers)
- Experience implementing HIPAA and SOC 2 compliance
- Experience operating in an HPC environment
- Bachelor's or Master's in Computer Science, Computer Engineering, or a related field