SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Platform Operations Engineer, Infrastucture

Viome - Bellevue, WA, United States - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Viome is seeking a Senior Platform Operations Engineer to lead the simplification and migration of its Azure-based service platform. The core mandate is to consolidate a complex, abstracted infrastructure stack onto a deliberately stable, Unix-based target platform: Linux VMs, systemd-supervised services, Apache, HAProxy, and Nginx routing. Key Responsibilities: Simplification & Migration: Plan and execute zero-downtime migrations of services from AKS/containers to VM-based hosting. Author systemd unit files, configure reverse proxies and TLS termination, and automate certificate management. Build and own deployment scripts (bash-based) that enforce immutable artifacts, commit traceability, and scripted rollback. Inventory and retire stale infrastructure—dormant deployments, unused DNS records, orphaned firewall rules. Maintain host hygiene: OS patching, runtime vendoring, log rotation, and centralized aggregation. Operate and migrate data stores (PostgreSQL, MySQL). Network & Infrastructure Operations: Troubleshoot workloads during transition, focusing on networking, load balancing, and core platform services. Administer Azure hub networks (Firewall rules, VPN gateways, VNet peering, DNS zones). Support the existing release process and operations until each service migrates. Maintain observability (OpenTelemetry, ELK, Grafana, uptime/cost monitoring). Coordinate cross-cloud dependencies with AWS (DNS, queues, egress allowlists). Provide L2+ operations support and manage runbooks. External Integrations & Security: Own integrations at risk during migration—e-commerce, subscription platforms, messaging providers, clinical/health-data partners. Ensure webhook delivery, signature verification, idempotency, and retry semantics. Raise security baseline: secrets management, webhook authentication, least-privilege access, data-retention hygiene. Required Qualifications: - 7+ years operating production Unix/Linux systems with deep systemd and process supervision knowledge - Strong shell scripting (bash) plus one additional language (Python/PHP) with a track record of building deploy/rollback tooling - Demonstrated reverse-engineering ability for undocumented production services - Message-queue literacy (Azure Service Bus, SQS, RabbitMQ) - PostgreSQL operations expertise (replicas, connection proxying, diagnostics) - Working proficiency with Kubernetes and cloud networking (private AKS, Istio-style ingress, Azure firewall/DNS) - Migration experience with zero-downtime cutover and tested rollback - Demonstrable record of reducing system surface area - Clear written communication for runbooks and cross-functional coordination Strongly Preferred: - Azure networking at landing-zone depth (hub-spoke VNets, Azure Firewall, Private Link/Private DNS) - External-dns, cert-manager, and policy engines (Kyverno) administration - ELK and OpenTelemetry pipeline operations - Webhook-heavy integration experience (Shopify, subscription billing, healthcare data exchanges) - Regulated/health-data environment experience (PHI handling, audit trails) - Terraform or equivalent IaC for cloud resources - AWS cross-cloud management (Route53, SQS, egress allowlists) - Proficiency with Claude or similar AI tools Success Metrics: - First 30 days: Trace and resolve a production issue end-to-end; produce a true-dependency inventory for one candidate service - First 90 days: Complete migration of 2–3 services with zero downtime; retire 20% of stale infrastructure

Similar roles