SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Mirantis - Hyderabad, India - In-office - posted 2026-09-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Mirantis is seeking a Senior Site Reliability Engineer to own both the developer platform and a customer-facing production region of a multi-tenant control plane for enterprise GPU infrastructure. You will build and run the environments, pipelines, and infrastructure tooling that engineering teams across the US, Europe, and APAC depend on to ship daily. This role spans development and production operations. You will make development, test, and pre-production clusters fast and reproducible, harden the Helm and CI/CD path from commit to release, and carry operational ownership—including on-call—for one of the smaller customer-facing production regions under real availability commitments. That production experience makes you the internal expert on how k0rdent AI is deployed and operated, consulted by teams running larger regions. Key responsibilities include: - Own the Kubernetes footprint across development, CI, pre-production, and one customer-facing production region (local kind clusters, shared dev/QA environments, multi-cluster/multi-region topologies) - Operate your production region against defined SLOs: capacity and upgrade planning, patching, backup/restore, disaster recovery drills, and on-call rotation participation - Lead incident response for your region: detection, mitigation, customer-impact assessment, root-cause analysis, and blameless postmortems - Build and maintain Helm charts and umbrella releases with versioning, values hygiene, and upgrade paths - Own CI/CD pipelines end-to-end: build, test, image publishing, chart packaging, release cutting, hotfix/backport flows - Automate environment bootstrap and seeding so any engineer can bring up a full stack with one command - Operate and troubleshoot the supporting stack: PostgreSQL, Temporal, Keycloak, API gateway, message broker, observability components - Build observability and diagnostics: metrics, dashboards, alerting, log and audit access for engineering and production - Consult with product teams and larger-region operators on deployment topology, GPU scheduling, RBAC, networking, failure modes; validate upgrade/migration procedures and hand over runbooks - Enforce security and tenant isolation: least-privilege access, secret handling, certificate/TLS lifecycle, image/dependency scanning, audit evidence for compliance - Drive infrastructure as code and repeatability—no snowflake environments, no undocumented manual steps - Mentor engineers on Kubernetes and operational practice; raise the team's bar through review and documentation Requirements: - 10+ years in DevOps, SRE, platform, or infrastructure engineering, including production ownership of customer-facing Kubernetes environments - Expert-level Kubernetes: workloads, networking, storage, RBAC, resource management, CRDs and operators, cluster upgrades—able to debug from kubectl and cluster internals - Proven incident response under SLA pressure: on-call rotations, escalation paths, postmortems, corrective action follow-through - Strong CI/CD engineering: pipelines as code, reproducible builds, artifact and release management (GitHub Actions or equivalent) - Solid scripting and automation ability; enough Go familiarity to read service code, trace failures, and file precise bugs - Experience running stateful supporting stack in Kubernetes: relational databases, identity providers, gateways, message brokers—including backup, restore, upgrade - Track record as technical consultant to other engineering teams: clear runbooks, design feedback, incident write-ups across global time zones (strong written English) - Must-have technical depth in several of: Kubernetes at scale, Cluster API, controllers/operators, Docker, Helm; GitHub Actions or equivalent CI/CD; Terraform/Ansible or equivalent plus GitOps (Argo CD, Flux); Keycloak and API gateway operation; PostgreSQL operations and migrations, streaming/message-broker platforms; Prometheus, Grafana, centralized logging, alerting tied to SLOs; AWS networking, IAM, load balancing, managed Kubernetes - Bachelor's degree in Computer Science & Engineering or related field, or 10+ years related experience Nice-to-have: k0s/k0rdent ecosystem experience, GPU infrastructure on Kubernetes, Temporal operations, bare-metal provisioning, multi-region topologies, service mesh, cross-cluster networking, Python for test harnesses, load/performance testing, policy enforcement (OPA/Kyverno), secret management, OpenTelemetry, distributed tracing, SLO/error-budget practice, compliance exposure (SOC 2, ISO 27001), CNCF open-source contributions.

Similar roles