SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure serving hundreds of financial institutions across 40 countries with institutional-grade APIs. The company is backed by $400 million in funding from top-tier investors including Portage Ventures, Spark Capital, and Y Combinator, with a team of 400+ globally distributed members.
You will lead Alpaca's incident operations function, building and scaling the team that commands the company's most critical incidents. This is a leadership role focused on establishing and continuously improving the incident response capability across EMEA and AMER regions, with oversight of a 24x7 follow-the-sun rotation.
Key responsibilities include:
**Building the Team & Command Structure**: Recruit and certify Incident Commanders, establish 24x7 coverage across APAC, EMEA, and AMER with warm handoffs at regional boundaries, and carry a rostered slot yourself. Develop a blameless review culture and keep the team sharp through game days, tabletop exercises, and simulations.
**Process Ownership & Maturity**: Drive severity model development with Risk on financial and regulatory materiality (critical in regulated brokerage where severity calls trigger reporting obligations). Own escalation paths, define thresholds for engineering leadership involvement, and maintain an accurate service catalogue with clear ownership.
**Bridge Operations**: Connect engineering teams (who need uninterrupted focus) with partner communications teams (who need continuous accurate updates). Run retrospectives with SRE, build post-incident packages with clear ownership and ticketing, and synthesize learnings for organization-wide distribution.
**KPI Ownership**: Define and track time-to-respond and time-to-mitigate metrics end-to-end, including data hygiene and definitions. Establish defensible baselines, move targets by severity, and track postmortem review overdue items by team.
**Automation & AI Integration**: Build incident operations as a documented, versioned, deployable product. Own the roadmap for AI workflows and agents that can set up incidents, assemble timelines, draft RCAs, and chase overdue updates—while maintaining clear boundaries on what requires human judgment.
You will own how well Alpaca responds to incidents, not the technical fix, partner communication, or reliability standards themselves. This deliberate separation of concerns is core to the role design.
**Requirements:**
Must-haves:
- Proven track record standing up (not just working within) an incident command or major-incident function, including ownership of severity models, roster building, and cross-team adoption
- 5+ years in production engineering, SRE, or technical operations with hands-on command of high-severity incidents
- Experience leading distributed teams across time zones and running 24x7 rotations
- Ability to influence engineers you don't manage and defend severity calls to stakeholders who disagree
- Track record building reliability metrics that people trust, with clear understanding of improving numbers vs. improving reality
- Disciplined scope management—can say "that's not ours" and route appropriately mid-outage without leaving gaps
- Strong writing and communication skills; ability to hold bridges calm under pressure and brief executives accurately without downplaying or dramatizing
- Understanding of FinTech and the trust stakes of API-driven financial platforms
- Experience using AI and agentic automation to reduce toil
Nice-to-haves:
- Formal incident command training (ITIL, Major Incident Management, crisis management)
- Experience running certification, game day, or drill programs with certified responder pools
- Familiarity with modern incident management and on-call platforms
- Track record building and maintaining service catalogues or ownership registries
- Experience converting incident follow-ups into funded roadmap work with program management or reliability functions
- Knowledge of incident reporting obligations in regulated financial services (DORA, Reg SCI, FINRA, or equivalent)
- Online securities trading, capital markets, or other regulated, market-hours-sensitive domain experience
- Experience deploying the same operating model across multiple regions or entities