SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
K Health is seeking a Senior DevOps Engineer to join its DevOps team and own the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role based in New York City (hybrid, 4 days/week in office) with participation in a daytime on-call rotation.
You will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will mentor junior engineers and collaborate closely with product and engineering teams across the company.
Key responsibilities include:
- Own the design, implementation, and evolution of GKE-based Kubernetes infrastructure across K Health and enterprise partner environments
- Build and maintain a Terraform modular infrastructure library with automated testing across GCP, Cloudflare, and AWS
- Architect and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment)
- Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others
- Implement and support security and compliance controls across infrastructure and the software supply chain—secrets management, pipeline secret detection, container scanning, SOC2, and HIPAA
- Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests
- Lead development of AI-powered operations tooling and agentic infrastructure
- Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts
- Mentor junior DevOps engineers and establish team-wide engineering standards
K Health is an AI-powered virtual care platform founded in 2016, headquartered in New York City, and backed by over $443.5M from leading investors. The company modernizes primary care by using AI to put humans first, offering clinical AI solutions for patients, provider-serving agentic solutions to reduce administrative burden, and AI-powered virtual clinics for health systems.
Requirements:
- 5+ years of experience in DevOps, platform engineering, or site reliability engineering
- Deep, hands-on experience with Kubernetes and the surrounding ecosystem (Helm, Helmfile, ArgoCD, Kyverno, cert-manager, NGINX Ingress)
- Extensive experience with Google Cloud Platform (GKE, Cloud SQL, Memorystore, Cloud Storage, IAM, Workload Identity)
- Strong Terraform expertise: modular architecture, multi-environment provisioning, and automated testing
- Advanced knowledge of GitLab CI/CD and GitOps practices
- Proficiency in Python and/or Go
Desired qualifications:
- Advanced Bash scripting skills
- Experience with secrets management solutions (Akeyless, HashiCorp Vault)
- Database administration experience across PostgreSQL, Redis, and MongoDB, including DR configuration and operational runbooks
- Experience with Datadog or equivalent observability platform (APM, infrastructure, log management)
- Experience with Cloudflare for DNS, CDN, and security rules management
- Demonstrated experience designing and executing disaster recovery programs
- Experience in highly regulated environments (SOC2, HIPAA)
- Excellent communication skills and ability to lead cross-functional infrastructure initiatives
- Demonstrated leadership experience, including mentoring junior engineers
- Experience with HPC or GPU cluster infrastructure (Slurm)
- Experience building or operating AI agents or agentic infrastructure
- Experience with microservices architecture and API gateway/reverse proxy patterns
- Experience with AWS