SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, Robinhood Command Center

Robinhood - Menlo Park, CA, United States - Hybrid - posted 2026-09-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 196,000 - 230,000 / annual

Robinhood is seeking a Staff Software Engineer to join the newly formed Robinhood Command Center (RCC), a reliability team serving as the front line for detecting, coordinating, and mitigating production incidents across the company. You will be part of the founding RCC team, helping define how Robinhood responds to and learns from incidents at scale. This is a highly visible role focused on incident leadership, operational excellence, and reliability tooling. You will own the processes and tools that enable fast, high-quality incident response, rather than owning product services or core infrastructure directly. Key responsibilities include: - Serve as a technical leader driving long-term reliability and observability strategy across Robinhood's infrastructure - Partner across engineering teams to raise the bar for operational excellence and incident response - Lead incident mitigation efforts by coordinating service owners and facilitating time-sensitive decisions like rollbacks and traffic shifts - Develop and maintain incident management processes and procedures to ensure timely resolution and minimize customer impact - Own incident discovery at the company level by defining and maintaining global dashboards and alerts tied to critical user journeys, availability, and business-impact metrics - Own and evolve incident response tooling and processes, including education, adoption, and measurement of MTTD/MTTR improvements - Drive post-incident governance and learning, defining standards for postmortems, SEV reviews, and follow-up tracking - Design and implement next-generation failure mitigation strategies that avoid full-region or full-datacenter failovers - Define and build frameworks to improve monitoring, alerting, and observability across hundreds of services and systems - Define and own the roadmap for bringing observability to critical user journeys for Robinhood's products - Deliver key insights and executive-level reporting to enable better business decisions around service quality and reliability - Act as a force multiplier through mentoring, technical influence, and contributions to hiring and engineering culture The role is based in Menlo Park, CA, with in-person attendance expected at least 3 days per week. Requirements: - 8+ years of software engineering experience, including significant experience operating production systems - 3+ years focused on reliability engineering, infrastructure, distributed systems, or production operations - Hands-on experience serving in incident leadership roles (e.g., IMOC, incident commander, primary oncall) - Strong communication and cross-functional collaboration skills, especially during high-severity incidents - Deep knowledge of systems reliability, observability frameworks, and fault-tolerant architecture design - Experience with multi-region or multi-cluster architectures, capacity planning, and failover strategies - Familiarity with modern observability stacks (e.g., OpenTelemetry, Prometheus, Grafana) - Demonstrated ability to drive measurable improvements in MTTD, MTTR, availability, or customer impact

Similar roles