SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer

Razorpay - Bengaluru, India - In-office - posted 2026-09-24

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Razorpay is India's leading full-stack fintech company, processing $180+ billion in annualized transactions and powering millions of businesses across India, Singapore, and Malaysia. The company offers a comprehensive platform spanning payments, payroll automation, fraud detection, and financial insights, backed by investors including GIC, Peak XV Partners, and Tiger Global. You will be one of the founding Site Reliability Engineers embedded with Razorpay's payment platform teams. Your mandate is to elevate payment flow reliability from three nines to four and five nines of availability. In payments, a failed request means a customer's money in limbo—you will define what reliability means, build the systems that enforce it, and set the standard for all future SREs. Key responsibilities: - Define SLIs, SLOs, and error budgets for critical payment flows (authorization, capture, refunds, settlements, webhooks) and establish them as shared language between product and platform teams. - Own the release lifecycle for payment services: design progressive rollout pipelines (canary, staged, feature-flagged), automated rollback triggers, and enforce sub-5-minute rollback capability as a launch-blocking requirement. - Carry the pager for payment-critical services, lead incident command during outages, and drive blameless postmortems with shipped action items. - Eliminate toil through software: build automation for failover, capacity management, load shedding, and degradation to prevent known failure classes from recurring. - Harden payment flows against distributed systems failure modes: retry storms, thundering herds, cascading failures, partial outages of banks and network partners, idempotency violations, and reconciliation gaps. - Run production readiness reviews for new payment services and hold the line on launch gates using error budget data, not opinion. - Design alerting that pages on customer-facing symptoms, not noise, and drive quarter-over-quarter improvements in mean time to detection and recovery. - Practice failure on purpose: organize game days, chaos experiments, and failure injection against payment-critical paths. Requirements: - 10+ years of engineering experience, with at least 5 years operating large-scale distributed systems in production (high QPS, multi-region, or systems where sub-1% error rates were business-critical). - Strong software engineering skills in at least one of Go, Java, or Python; you have built tools and services, not just configured them. - Deep understanding of distributed systems failure modes and containment patterns: circuit breakers, backpressure, bulkheading, graceful degradation, idempotency. - Solid fundamentals in Linux internals, networking, and databases under load (replication, failover, connection pool exhaustion, lock contention). - Hands-on experience designing or significantly improving deployment pipelines: canary analysis, automated rollback, feature flags. - Genuine on-call ownership: you have carried a pager for systems that mattered, led incidents, and can walk through a specific outage you handled and the changes you made afterward. - Fluency with modern observability (metrics, tracing, structured logging; e.g., Prometheus, Grafana, OpenTelemetry, Datadog, Coralogix, Clickhouse or similar) and experience reducing alert noise. - Experience defining SLOs and error budget policies from scratch; contributions to reliability tooling (open source or internal) that other teams adopted. - Judgment and communication skills to tell product teams "not yet" with data, and pragmatism to help them get to "yes" quickly. Nice to have: - Experience in payments, fintech, banking, trading, or another domain where correctness and money are coupled (transactional consistency, exactly-once semantics, reconciliation). - Experience with Kubernetes at scale, service mesh, and traffic management.

Similar roles