SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer/Cloud Platform Engineer

Skyflow - India - In-office - posted 2026-08-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Skyflow is seeking a Senior Platform Engineer to design and operate the cloud infrastructure powering a multi-tenant, security-sensitive B2B SaaS product serving Fortune 500 enterprises. The role combines hands-on software engineering with deep operational ownership across AWS and GCP, spanning both multi-tenant and single-tenant (BYOC) customer deployments. You will architect and build automation for end-to-end infrastructure provisioning, upgrades, scaling, and decommissioning across multiple cloud providers and deployment models. This is primarily a software engineering role—you'll write production-grade Go and Python services, CLIs, and infrastructure-as-code (Terraform/OpenTofu, Helm, ArgoCD) that transform manual operations into self-service platform capabilities for other engineering teams. Key responsibilities include: designing provisioning pipelines and internal tooling; owning and evolving the IaC stack; operating core platform services (Kubernetes, Istio, Aerospike, PostgreSQL, Kafka, GPU-backed inference); building observability and alerting to catch failures before customers notice; participating in on-call rotations and leading incident response; and partnering with security/compliance to embed guardrails into the platform by default. You'll bring 8+ years of production platform engineering, infrastructure, DevOps, or SRE experience with strong software engineering skills in Go and/or Python. You need deep hands-on Kubernetes expertise in production, solid experience with at least one major public cloud (AWS or GCP preferred), and a track record of building tools or platforms that other engineers use. You're comfortable in B2B enterprise environments where infrastructure changes carry compliance and contractual weight, have strong incident-response instincts, and bias toward root-causing and automating away recurring problems. Nice-to-have skills include service mesh (Istio/Envoy), GitOps workflows (ArgoCD/Flux), policy-as-code (OPA), experience with stateful systems (distributed databases, message queues), security-sensitive/regulated environments (PCI, SOC 2, HIPAA), cost optimization at fleet scale, and open-source infrastructure contributions.

Similar roles