SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, Reliability

Metropolis Technologies - Bengaluru, Karnataka, India - In-office - posted 2026-07-31

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Metropolis Technologies is building AI-powered infrastructure for the Recognition Economy, transforming parking and expanding into retail and hospitality. They are seeking a Staff Software Engineer focused on Reliability to own the reliability posture across the entire platform and drive comprehensive practices ensuring system availability, resilience, and observability for mission-critical mobility infrastructure. In this role, you will be the technical owner of reliability, architecting failover systems, implementing chaos engineering, and improving observability foundations to maintain 99.9%+ uptime as the company scales to new markets. You will join the Product Foundations team and work alongside highly technical teams to influence architecture decisions and establish company-wide reliability standards. Key responsibilities include: owning overall reliability posture and establishing practices/metrics for 99.9%+ uptime; designing automatic failover mechanisms for critical external dependencies (Twilio, Stripe) with circuit breakers and retry policies; architecting active-passive or active-active regional deployment strategies with database replication and automated failover; establishing comprehensive monitoring using Datadog for APM, logs, and metrics; implementing synthetic monitoring, SLO-based alerting, and on-call rotation; owning incident management processes including post-mortem culture and MTTR reduction; driving adoption of resilience patterns across services including health checks, graceful degradation, and feature flags; and building local mirrors for critical dependencies with artifact caching and vulnerability scanning. Required qualifications: 8+ years of engineering experience including software engineering, reliability engineering, SRE practices, or production operations at scale; expert-level reliability engineering skills with hands-on experience in multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery; production observability expertise with deep experience implementing monitoring, alerting, tracing, and logging systems at scale using Datadog or similar APM platforms; strong systems thinking and ability to design resilient distributed systems; database and data systems knowledge including replication strategies, backup/restore, and experience with relational and NoSQL databases; cloud platform expertise with production experience on AWS including multi-region deployments; experience with AI-powered development tools such as Claude Code or GitHub Copilot; excellent technical communication and ability to influence across teams; and expert-level Java and/or Scala proficiency with strong understanding of JVM performance and concurrency. Plus factors include Scala experience and SRE/Reliability Engineering experience at companies known for operational excellence such as Google, Amazon, or Netflix.

Similar roles