SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer - Robinhood Command Center

Robinhood - Menlo Park, CA, United States - Hybrid - posted 2026-04-23

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Robinhood is seeking a Senior Software Engineer to join the newly formed Robinhood Command Center (RCC), a reliability team serving as the front line for detecting, coordinating, and mitigating production incidents across the company. This is a founding team role focused on incident leadership, operational excellence, and reliability tooling. You will serve as a senior technical leader driving long-term reliability and observability strategy across Robinhood's infrastructure. Key responsibilities include: leading incident mitigation efforts by coordinating service owners and facilitating time-sensitive decisions like rollbacks and traffic shifts; developing and maintaining incident management processes to ensure timely resolution and minimize customer impact; owning incident discovery at the company level by defining and maintaining global dashboards and alerts tied to critical user journeys, availability, and business-impact metrics; owning and evolving incident response tooling and processes, including education, adoption, and measurement of MTTD/MTTR improvements; driving post-incident governance and learning through postmortem standards and SEV reviews; designing and implementing next-generation failure mitigation strategies; defining and building frameworks to improve monitoring, alerting, and observability across hundreds of services; and delivering key insights and executive-level reporting on service quality and reliability. You will partner closely across many different types of engineers to raise the bar for operational excellence and act as a force multiplier through mentoring, technical influence, and contributions to hiring and engineering culture. Required qualifications: 5+ years of software engineering experience with significant production systems operation; 2+ years focused on reliability engineering, infrastructure, distributed systems, or production operations; hands-on experience in incident leadership roles (IMOC, incident commander, primary oncall); strong communication and cross-functional collaboration skills, especially during high-severity incidents; deep knowledge of systems reliability, observability frameworks, and fault-tolerant architecture design; experience with multi-region or multi-cluster architectures, capacity planning, and failover strategies; familiarity with modern observability stacks (OpenTelemetry, Prometheus, Grafana); and demonstrated ability to drive measurable improvements in MTTD, MTTR, availability, or customer impact. The role is based in Menlo Park with in-person attendance expected at least 3 days per week.

Similar roles