SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure serving hundreds of financial institutions across 40 countries. The company provides institutional-grade APIs for stocks, ETFs, options, crypto, fixed income, and 24/5 trading, supporting over 10 million brokerage accounts. Backed by $400 million in funding from top-tier investors including Spark Capital, Tribe Capital, and Y Combinator, Alpaca operates a globally distributed team of 400+ members.
As Incident Operations Commander, you will serve as the on-duty commander for Alpaca's most critical incidents, directing cross-functional response to restore service quickly while keeping the right people engaged and informed. Your role is not to fix the outage, but to make the response reliable: ensuring correct severity classification, assembling the right engineers, enabling mitigation without stalls, keeping leaders informed, and ensuring follow-up work survives the incident call.
Key responsibilities include:
• Command incidents end-to-end from declaration through mitigation, keeping responders focused on stopping customer and partner impact as fast as possible. Run the bridge, protect responders from distractions, and identify stalls in real time.
• Classify and hold the line on severity at declaration, re-checking as facts arrive. Work with Risk on financial and regulatory materiality, but own the severity call.
• Engage the right people rapidly by identifying owning teams by service, symptom, and blast radius. Page them immediately and expand the responder set when needed. Escalate unanswered pages and bring in leaders who must make business calls (feature flags, traffic shedding, failover, freeze-or-ship decisions).
• Hold the bridge and protect the people fixing it. Be the single point of contact for stakeholder, partner, and executive questions. Serve as the authoritative source for the partner communications team on impact, severity, and timing.
• Run follow-the-sun handoffs across regions with warm, high-fidelity transitions: current impact and severity, mitigation path, next actions, who is in the room, outstanding decisions, and what must not be dropped.
• Close the loop on schedule. Maintain a timeline of facts as the incident runs (critical for regulated business records). Once mitigated, ensure a blameless retrospective is scheduled with a named owner and timebox. Record where the cause sits to determine which follow-up items are mandatory. Every action item needs a real ticket, one accountable owner, priority, and category, delivered within the agreed service level.
• Automate coordination away. Identify manual prompts—updates due, unanswered questions, uncontacted partners—as candidates for automation. Work toward AI handling routine coordination so commanders can focus on judgment.
You will work a regional coverage window as part of a global 24x7 Incident Commander roster.
REQUIREMENTS (Must-Haves):
• 4+ years commanding or co-commanding high-severity incidents in a production engineering, SRE, or technical operations environment
• Ability to direct technical responders under pressure without being the person writing the fix
• Proven ability to make and defend crisp severity and escalation decisions; take charge without waiting to be asked; wake senior people at 03:00, interrupt executives, and redirect experienced engineers with confidence
• Can read a dashboard and judge for yourself whether impact has actually stopped
• Clear communication with engineers, executives, and partner-facing stakeholders; understand the difference between briefing comms and speaking for the company
• Comfortable holding other teams accountable in the moment across reporting lines that are not yours, without creating friction
• Thrive in a follow-the-sun model with clean cross-region handoffs
• Understand FinTech concepts and the trust stakes of API-driven financial platforms
• Use AI tools and agentic automation to reduce manual toil and speed up response
• Willing to work a regional coverage window as part of a global 24x7 roster
NICE-TO-HAVES:
• Formal incident command training (ITIL, Major Incident Management, or crisis management)
• Experience with modern incident management and on-call platforms
• Have written severity rubrics, decision trees, escalation matrices, runbooks, or incident playbooks
• Commanded in game days, tabletop exercises, or incident simulations, not only in production
• Partnered with problem management or reliability programme functions to roadmap incident follow-ups
• Online securities trading, capital markets experience, or other regulated, market-hours-sensitive domain background