SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 203,500 - 248,500 / annual
Metropolis Technologies is building AI-powered infrastructure for the Recognition Economy, transforming parking and expanding into retail and hospitality. As Staff Software Engineer focused on Reliability, you will own the reliability posture across the entire Metropolis platform, architecting and implementing systems that ensure 99.9%+ uptime for mission-critical mobility infrastructure.
Key responsibilities include:
- Own overall reliability posture, establishing practices, metrics, and systems ensuring 99.9%+ uptime across all services
- Design and implement automatic failover mechanisms for critical external dependencies (Twilio, Stripe) with circuit breakers, retry policies, and degraded mode operations
- Architect active-passive or active-active regional deployment strategies with database replication, automated failover, and DNS-based traffic routing including disaster recovery planning
- Establish comprehensive monitoring using Datadog for APM, logs, and metrics correlation
- Implement synthetic monitoring, SLO-based alerting, on-call rotation, and escalation policies while building service health dashboards showing customer impact
- Own incident management process including workflows, tooling, post-mortem culture, runbook automation, and MTTR reduction initiatives
- Drive adoption of resilience patterns across services including health checks, graceful degradation, feature flags, rate limiting, backpressure mechanisms, and chaos engineering
- Build and maintain local mirrors for critical dependencies with artifact caching, dependency pinning, and vulnerability scanning
You will join the Product Foundations team, playing a key role in building foundational infrastructure powering the future of mobility commerce. The role requires 4 days in-office weekly to foster collaboration and innovation.
Tech stack: TypeScript, React, Scala (principal), Java (limited), MySQL, PostgreSQL, Snowflake, AWS, GitHub, Datadog.
REQUIREMENTS:
- 8+ years of engineering experience including software engineering, reliability engineering, SRE practices, or production operations at scale
- Expert-level reliability engineering skills with hands-on experience in multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery
- Production observability expertise with deep experience implementing monitoring, alerting, tracing, and logging systems at scale—specifically Datadog or similar APM platforms in high-load environments
- Strong systems thinking with proven ability to design resilient distributed systems handling failures, network partitions, and external dependency outages
- Database and data systems knowledge including replication strategies, backup/restore procedures, connection pooling, query optimization, and experience with relational and NoSQL databases
- Cloud platform expertise with production experience operating systems on AWS including multi-region deployments, load balancing, and DNS-based failover
- Experience with AI-powered development tools such as Claude Code, GitHub Copilot, or similar agentic coding tools—context engineering in particular
- Excellent technical communication with ability to influence technical decisions across teams, document complex systems, conduct post-mortems, and establish reliability standards organization-wide
- Expert-level Java and/or Scala proficiency with strong understanding of JVM performance, concurrency, and operational characteristics
PLUS (not required):
- Scala experience
- SRE or Reliability Engineering experience at companies known for operational excellence (Google, Amazon, Netflix) or high-growth startups where you built reliability practices from the ground up
- Incident response leadership including building incident management processes, conducting blameless post-mortems, and driving MTTR reduction
- Chaos engineering experience with tools like Chaos Monkey, Gremlin, or similar
- Performance optimization experience with profiling, benchmarking, capacity planning, and system tuning at hyperscale
- Open source contributions or technical blog writing demonstrating expertise in reliability engineering, distributed systems, or production operations