SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Robinhood is seeking a Senior Software Engineer to join the newly formed Robinhood Command Center (RCC), a reliability team serving as the front line for detecting, coordinating, and mitigating production incidents across the company. This is a founding team role focused on incident leadership, operational excellence, and reliability tooling.
You will serve as a senior technical leader driving long-term reliability and observability strategy across Robinhood's infrastructure. Key responsibilities include: leading incident mitigation efforts by coordinating service owners and facilitating time-sensitive decisions like rollbacks and traffic shifts; developing and maintaining incident management processes to ensure timely resolution and minimize customer impact; owning incident discovery at the company level by defining and maintaining global dashboards and alerts tied to critical user journeys, availability, and business-impact metrics; owning and evolving incident response tooling and processes, including education, adoption, and measurement of MTTD/MTTR improvements; driving post-incident governance and learning through postmortem standards and SEV reviews; designing and implementing next-generation failure mitigation strategies; defining and building frameworks to improve monitoring, alerting, and observability across hundreds of services; and delivering key insights and executive-level reporting on service quality and reliability.
You will partner closely across many different types of engineers to raise the bar for operational excellence and act as a force multiplier through mentoring, technical influence, and contributions to hiring and engineering culture.
Required qualifications: 5+ years of software engineering experience with significant production systems operation; 2+ years focused on reliability engineering, infrastructure, distributed systems, or production operations; hands-on experience in incident leadership roles (IMOC, incident commander, primary oncall); strong communication and cross-functional collaboration skills, especially during high-severity incidents; deep knowledge of systems reliability, observability frameworks, and fault-tolerant architecture design; experience with multi-region or multi-cluster architectures, capacity planning, and failover strategies; familiarity with modern observability stacks (OpenTelemetry, Prometheus, Grafana); and demonstrated ability to drive measurable improvements in MTTD, MTTR, availability, or customer impact.
The role is based in Menlo Park with in-person attendance expected at least 3 days per week.