SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer III

Vida Health - Remote - Remote - posted 2026-09-30

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Vida Health is seeking a Site Reliability Engineer III to be the company's first dedicated SRE hire. You'll join the Enablement Team, which owns the platform and tooling that engineering teams build on, reporting to the Engineering Manager and working closely with the Lead Engineer who will mentor you. Vida operates a virtual, personalized obesity care platform on GCP with a production GKE cluster hosting ~50 workloads (Django applications, Airflow jobs). The data layer includes Cloud SQL (MySQL and PostgreSQL), Redis, and Firestore. Infrastructure is defined across two Terraform repositories that have grown organically with inconsistent patterns. In this role, you'll modernize and scale infrastructure to support a wave of new enterprise contracts launching January 1. Key responsibilities include: • Consolidate Terraform repositories into a clean, well-documented structure with consistent patterns for state management, module structure, code review, and CI checks • Normalize environments and improve build/deploy automation in GitHub Actions; add drift detection and alerting • Apply overdue patches and upgrades across Cloud SQL databases and application runtimes • Right-size compute and database workloads for growth, including connection pooling and scaling improvements for high-traffic services • Evaluate Kubernetes architecture for multi-cluster readiness as the company scales • Improve monitoring and observability in Datadog and Cloud Monitoring to catch issues before they become incidents • Design observability access for contractors and external partners while protecting PHI • Retire legacy infrastructure and tooling • Build repeatable operational processes: runbooks, on-call rotation, escalation documentation • Support infrastructure readiness for January 1 enterprise launches This is a fully remote role with no time zone restrictions. You'll help shape what SRE looks like at Vida going forward. REQUIREMENTS: • Bachelor's degree minimum • 5+ years of experience in SRE, DevOps, or infrastructure engineering with real ownership of production systems • Deep hands-on Terraform experience, including structuring modules and managing state across environments • Strong working knowledge of GCP: GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking, load balancing, cost management • Production Kubernetes experience: autoscaling, resource management, architectural judgment • Hands-on experience building monitoring, alerting, and dashboards with Datadog or Cloud Monitoring • Proficiency in Python for tooling and automation • Comfortable working across multiple teams and explaining infrastructure decisions to non-specialists PREFERRED: • Experience as an early or first SRE hire • Experience refactoring or consolidating a large, organically grown Terraform codebase • Experience improving observability from a less mature baseline • Experience in HIPAA-regulated or compliance-driven environments • CI/CD experience with GitHub Actions • Experience running Django applications or Airflow in production on Kubernetes • Experience designing or migrating to multi-cluster Kubernetes architectures

Similar roles