SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer - India

JumpCloud - Remote - Remote - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

JumpCloud is seeking a Site Reliability Engineer (Software Engineer 3) to join the Infrastructure & Reliability Engineering team. This is an engineer-first role focused on designing automation, building observability frameworks, and reducing operational toil through code rather than traditional operations work. You will be responsible for designing, deploying, and maintaining the reliability, availability, and performance of critical JumpCloud systems and APIs across AWS and GCP. Key responsibilities include: - Design and maintain SLIs, SLOs, and error budgets in partnership with application teams - Build end-to-end observability across microservices using tools like Datadog, implementing monitoring across Golden Signals (Latency, Traffic, Errors, Saturation) - Participate in on-call rotations, incident response, and blameless post-incident reviews - Manage and operationalize production Kubernetes (EKS) clusters using GitOps workflows (Argo CD, Kargo) - Provision and secure multi-cloud infrastructure using modular Terraform - Develop and maintain Disaster Recovery dashboards, runbooks, multi-region failover automation, and validation tests aligned with RTO/RPO targets - Eliminate operational toil by writing production-grade Python or Go automation tools - Leverage AI-assisted development tools to accelerate scripting and incident triage JumpCloud is an AI-powered unified IT management platform consolidating identity, device, and access management. The company operates in a fast-paced SaaS environment with teams across 15+ countries. The role includes on-call shift expectations and requires fluent English for interviews and internal communication. REQUIREMENTS: - 5+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical systems - Hands-on proficiency in Python or Go for SRE tools, custom automation, and cloud integrations - Production experience with Kubernetes cluster operations, container orchestration, and GitOps pipelines (Argo CD) - Solid experience writing, maintaining, and modularizing Terraform configurations - Direct experience operating cloud workloads on AWS (EKS, IAM, VPC networking, Route53, ALB/NLB) or GCP - Practical experience with FinOps: cost-allocation tagging, resource right-sizing, and FinOps dashboards - Experience building DR dashboards, running failover drills, and configuring monitoring tools - Practical experience with Datadog (or similar), PagerDuty, alerting hygiene, and SLI/SLO frameworks - Solid operational experience configuring and troubleshooting production service meshes (Istio or similar) and managing high-availability proxy solutions (HAProxy, NGINX) - Strong troubleshooting skills, effective collaboration, and track record of driving operational efficiency through code PREFERRED QUALIFICATIONS: - Experience with CI/CD tools such as GitHub Actions or GitLab Pipelines - Basic understanding of chaos engineering principles or testing resilience - Familiarity with secrets management tools (HashiCorp Vault, AWS Secrets Manager, External Secrets Operator) - Basic knowledge of DevSecOps tools and infrastructure-as-code vulnerability scanning

Similar roles