SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer/Cloud Platform Engineer - Operations (PST Timezone)

Skyflow - Remote - Remote - posted 2026-08-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Skyflow is seeking a Senior Platform Engineer to design, build, and operate the cloud infrastructure powering a multi-tenant, security-sensitive B2B SaaS product serving Fortune 500 enterprises. The role combines hands-on software engineering with deep operational ownership across AWS and GCP, supporting both multi-tenant and single-tenant (BYOC) deployments. You will spend most of your time writing production-grade Go and Python to automate infrastructure work—provisioning pipelines, internal CLIs, self-service tooling, and automation frameworks that eliminate manual toil. You'll own the Infrastructure-as-Code stack (Terraform/OpenTofu, Helm, ArgoCD), design and manage Kubernetes clusters at scale, and operate core platform services including service mesh (Istio), data stores (Aerospike, PostgreSQL), messaging (Kafka), and GPU-backed inference workloads. Key responsibilities include: designing end-to-end automation for environment provisioning and lifecycle management across multiple cloud providers and deployment models; building observability and alerting to catch failures before customers notice; participating in on-call rotations and leading incident response with root-cause analysis; partnering with security and compliance teams to embed guardrails (secrets management, access control, audit logging) into the platform; and continuously identifying and eliminating manual, repetitive infrastructure work. You'll provide operational support aligned with US time zones, ensuring system reliability and availability for enterprise customers with strict uptime SLAs and compliance requirements (PCI, SOC 2, HIPAA). Success is measured by reducing manual toil across the organization, not by ticket volume. Required: 4+ years in platform engineering, infrastructure engineering, DevOps, or SRE roles with real production cloud ownership; strong software engineering skills in Go and/or Python; deep hands-on Kubernetes experience in production; solid experience with AWS or GCP (both is a plus); track record of building tools and platforms for other engineers; comfort in B2B enterprise environments with compliance and security constraints; strong incident-response and debugging skills; bias toward root-causing and automating recurring problems. Nice to have: service mesh (Istio/Envoy), GitOps (ArgoCD/Flux), policy-as-code (OPA), experience with stateful systems (distributed databases, message queues, GPU workloads), security-sensitive/regulated environments, cost optimization and fleet-scale capacity planning, open-source infrastructure contributions, or internal developer platform (IDP) experience.

Similar roles