SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
JumpCloud is seeking a Senior Site Reliability Engineer to architect, scale, and continuously improve the reliability, availability, and performance of its multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
Key responsibilities include:
• Architect and maintain Disaster Recovery processes, multi-region failover automation, and business continuity strategies to meet strict RTO and RPO objectives.
• Lead the design and enforcement of SLIs, SLOs, and Error Budget frameworks across multi-disciplinary engineering teams.
• Drive end-to-end observability strategy using Datadog, implementing Golden Signals monitoring to reduce MTTD/MTTR and eliminate alert fatigue.
• Lead on-call escalation, major incident management, and enforce 99.99% availability SLAs.
• Facilitate blameless post-incident reviews and execute systemic root-cause remediations.
• Architect, manage, and scale production Kubernetes (EKS) clusters with advanced GitOps workflows (Argo CD, Kargo).
• Design and maintain enterprise-grade Infrastructure-as-Code using Terraform across multi-account, multi-region cloud environments.
• Design and build interactive FinOps and cost-optimization dashboards for engineering and leadership teams.
• Eliminate operational toil by writing production-grade Python or Go tooling, platform automation, and custom integrations.
• Champion AI-assisted software development workflows (Cursor, Claude Code, GitHub Copilot) for automation and incident triage.
• Author operational runbooks, architecture decision records, and mentor mid-level/junior engineers.
JumpCloud is an AI-powered unified IT management platform consolidating identity, device, and access management. The role is fully remote within India, with expectations for on-call shift participation. English fluency is required.
REQUIREMENTS:
• 8+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical, highly available distributed systems.
• Bachelor's degree in Computer Science, Software Engineering, or equivalent technical discipline.
• Advanced Python/Go capabilities for writing internal SRE platforms, tools, and API integrations.
• Deep Kubernetes expertise: hands-on production EKS/GKE cluster lifecycles, ingress/egress, networking, RBAC, and GitOps tooling (Argo CD).
• Advanced IaC and AWS/GCP proficiency: deep Terraform expertise (module architecture, state management refactoring) across complex multi-account AWS environments (IAM, VPCs, Transit Gateway, ALB/NLB, Route53).
• FinOps and cost optimization leadership: demonstrated experience driving cloud cost-efficiency strategies, resource right-sizing, cost-allocation tagging, workload optimization, and building FinOps dashboards.
• Disaster Recovery and high availability: proven background designing and testing multi-region DR architectures, automating failover systems, and monitoring recovery health.
• Observability and reliability architecture: track record defining SLI/SLOs, managing PagerDuty schedules, and optimizing production observability platforms.
• Experience designing and operating enterprise service meshes (Istio, Linkerd, or similar) and production ingress/proxy systems (HAProxy, NGINX, or similar).
• Technical mentorship: demonstrated ability to lead technical discussions, write architectural design docs/RFCs, and mentor engineering peers.
• Strong problem-solving, communication, and collaboration skills with passion for solving complex distributed systems challenges at scale.
PREFERRED QUALIFICATIONS:
• Basic understanding of chaos engineering principles or testing resilience in staging/production.
• Experience with secrets management architectures (Vault, AWS Secrets Manager, External Secrets Operator, Cert-Manager).
• Background in DevSecOps practices, service meshes (Istio), and automated vulnerability remediation within cloud infrastructure code.
• Background supporting identity services, IAM, enterprise directory platforms, or security-focused SaaS solutions.