SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Heidi Health is building AI-powered clinical tools to reduce administrative burden on healthcare providers. The company has achieved significant scale—supporting 2.5 million patient sessions weekly across 190+ countries—and is now expanding its product suite. This role sits within the model team, responsible for the infrastructure that trains, deploys, and operates the AI models powering Heidi's products.
You will own the end-to-end infrastructure for model serving, GPU cluster management, and the systems supporting training and evaluation. Your responsibilities include:
- Building and operating production model-serving infrastructure with repeatable deployment pipelines, request routing, autoscaling, fallback paths, and controlled rollouts across regions.
- Managing GPU clusters and workload scheduling to optimize resource allocation across inference, training, and evaluation while maintaining service responsiveness.
- Profiling and improving inference performance (latency, throughput, memory efficiency) through techniques like batching, KV-cache management, quantization, and speculative decoding.
- Supporting distributed training and model iteration by providing reliable job launching, checkpoint management, and validated model promotion to serving.
- Building observability and incident traceability, connecting requests to specific models, configurations, and workers; creating dashboards and alerts for latency, queueing, errors, GPU health, and workload performance.
- Owning production reliability through service objective definition, incident investigation across application/inference/GPU/network layers, and building recovery procedures and runbooks.
- Making compute costs actionable by tracking GPU usage, idle capacity, and inference cost by model and workload to guide deployment decisions.
- Building a self-service platform for the model team to provision, configure, benchmark, and release models independently across ASR, note generation, Evidence, and Dictate products.
You will partner closely with researchers and platform engineers, taking systems from initial design through rollout, incidents, and ongoing improvement.
REQUIREMENTS:
- Strong engineering foundation with hands-on AI infrastructure experience: at least 1 year building and operating infrastructure for large language models, including model deployment, inference serving, or distributed training. You must be able to design systems, write code, and own them in production.
- Production model deployment experience with LLMs or demanding ML workloads using engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or comparable systems.
- GPU cluster and orchestration experience managing GPU workloads using Kubernetes, Slurm, or equivalent platforms, with practical experience in scheduling, resource allocation, capacity planning, and failure recovery.
- Performance debugging skills using traces, metrics, and profiling tools to identify compute, memory, communication, and scheduling bottlenecks. Understanding of how batch size, context length, precision, and multi-GPU execution affect performance and cost.
- Strong software and systems skills: proficiency in Python and comfort with backend or systems development in Go, C++, Rust, or comparable languages. Practical experience with Linux, containers, deployment automation, and distributed services.
- Operational ownership: you've owned production incidents, built useful monitoring, made releases recoverable, and can explain design trade-offs while working effectively with researchers, product engineers, and infrastructure partners.
NICE TO HAVE:
- Experience with distributed training frameworks (PyTorch FSDP, Megatron) or infrastructure for reinforcement learning and rollout generation.
- Experience tuning inference engines, serving MoE models, or implementing quantization, speculative decoding, and prefill/decode disaggregation.
- Familiarity with GPU interconnects, NCCL, RDMA, topology-aware scheduling, or diagnosing multi-node communication problems.
- CUDA or Triton kernel development, contributions to AI infrastructure projects, or experience building cluster operators and scheduling integrations.
- Experience operating infrastructure across multiple regions or providers, particularly for healthcare or other sensitive production workloads.