SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 145,000 - 175,000 / annual
Crusoe is a vertically integrated AI infrastructure company building the full stack from energy to compute to power the world's most ambitious AI workloads. They are seeking a Senior Incident Manager to lead incident response and drive continuous improvement in reliability and customer trust.
In this role, you will own end-to-end incident management for high-visibility technical incidents and enterprise customer escalations. You'll coordinate response efforts across teams (without troubleshooting the technical issue yourself), act as a bridge between Customer Success and engineering/product teams, and own all customer-facing communications including status page updates and root cause analyses (RCAs). You'll maintain incident metrics and reporting to track performance, ensure corrective actions are documented and tracked, and develop training materials and knowledge base articles. The role includes participation in on-call rotation (10am–10pm, 5 days per week).
Key responsibilities include: leading incident response coordination across teams; owning customer communication and status page decisions; producing customer-facing RCAs that translate technical data into clear narratives for enterprise customers and executives; documenting and tracking corrective actions; maintaining incident metrics and reporting; designing incident response strategies and self-serve support processes as the function scales; and participating in on-call coverage.
This is a compelling opportunity to shape incident management processes at a high-growth company. The team has scaled from 2 people to 5, with continued growth ahead. You'll have exposure to cutting-edge AI infrastructure across the full stack: compute, networking (SDN), storage (VAST), Kubernetes, and Slurm. The role offers a builder mentality—the chance to create and refine processes, leave a lasting imprint, and grow alongside the team.
REQUIREMENTS:
- Background on an enterprise customer support team, or direct experience working with enterprise customers
- Strong written communication—able to turn technical, ambiguous incident detail into clear, structured narratives for a non-technical audience
- Understands the nuances of what can and cannot be communicated to customers, including what needs to be documented in writing
- Strong judgment on communication level—knowing what to filter out when reporting bad news, since each incident case is different
- Sound judgment on transparency—knowing when to stay high-level vs. go into detail
- Logical, structured thinker with a solid grounding in incident management process and best practices
- Comfortable working with data to track and report on trends, not just handling incidents in the moment
- Working familiarity with Linux, virtualization, and Kubernetes—comfortable enough to engage credibly with infrastructure and technical responders during an incident, without needing to be the one troubleshooting
BONUS:
- Experience working with customers in a startup environment
- Background in cloud infrastructure or AI-focused environments
- Understanding of virtualization and orchestration/scheduling technologies (Slurm or Kubernetes)
- Hands-on experience with Linux, virtualization, or Kubernetes
- Understanding of the TCP/IP stack
- Understanding of Infrastructure-as-Code (IaC) practices