SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer - India

JumpCloud - Remote - Remote - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

JumpCloud is seeking a Senior Site Reliability Engineer to architect, scale, and continuously improve the reliability, availability, and performance of its multi-region microservices, APIs, and authentication infrastructure on AWS/GCP. Key responsibilities include: • Architect and maintain Disaster Recovery processes, multi-region failover automation, and business continuity strategies to meet strict RTO and RPO objectives. • Lead the design and enforcement of SLIs, SLOs, and Error Budget frameworks across multi-disciplinary engineering teams. • Drive end-to-end observability strategy using Datadog, implementing Golden Signals monitoring to reduce MTTD/MTTR and eliminate alert fatigue. • Lead on-call escalation, major incident management, and enforce 99.99% availability SLAs. • Facilitate blameless post-incident reviews and execute systemic root-cause remediations. • Architect, manage, and scale production Kubernetes (EKS) clusters with advanced GitOps workflows (Argo CD, Kargo). • Design and maintain enterprise-grade Infrastructure-as-Code using Terraform across multi-account, multi-region cloud environments. • Design and build interactive FinOps and cost-optimization dashboards for engineering and leadership teams. • Eliminate operational toil by writing production-grade Python or Go tooling, platform automation, and custom integrations. • Champion AI-assisted software development workflows (Cursor, Claude Code, GitHub Copilot) for automation and incident triage. • Author operational runbooks, architecture decision records, and mentor mid-level/junior engineers. JumpCloud is an AI-powered unified IT management platform consolidating identity, device, and access management. The role is fully remote within India, with expectations for on-call shift participation. English fluency is required. REQUIREMENTS: • 8+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical, highly available distributed systems. • Bachelor's degree in Computer Science, Software Engineering, or equivalent technical discipline. • Advanced Python/Go capabilities for writing internal SRE platforms, tools, and API integrations. • Deep Kubernetes expertise: hands-on production EKS/GKE cluster lifecycles, ingress/egress, networking, RBAC, and GitOps tooling (Argo CD). • Advanced IaC and AWS/GCP proficiency: deep Terraform expertise (module architecture, state management refactoring) across complex multi-account AWS environments (IAM, VPCs, Transit Gateway, ALB/NLB, Route53). • FinOps and cost optimization leadership: demonstrated experience driving cloud cost-efficiency strategies, resource right-sizing, cost-allocation tagging, workload optimization, and building FinOps dashboards. • Disaster Recovery and high availability: proven background designing and testing multi-region DR architectures, automating failover systems, and monitoring recovery health. • Observability and reliability architecture: track record defining SLI/SLOs, managing PagerDuty schedules, and optimizing production observability platforms. • Experience designing and operating enterprise service meshes (Istio, Linkerd, or similar) and production ingress/proxy systems (HAProxy, NGINX, or similar). • Technical mentorship: demonstrated ability to lead technical discussions, write architectural design docs/RFCs, and mentor engineering peers. • Strong problem-solving, communication, and collaboration skills with passion for solving complex distributed systems challenges at scale. PREFERRED QUALIFICATIONS: • Basic understanding of chaos engineering principles or testing resilience in staging/production. • Experience with secrets management architectures (Vault, AWS Secrets Manager, External Secrets Operator, Cert-Manager). • Background in DevSecOps practices, service meshes (Istio), and automated vulnerability remediation within cloud infrastructure code. • Background supporting identity services, IAM, enterprise directory platforms, or security-focused SaaS solutions.

Similar roles