SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
xAI is seeking a NOC Technician to staff the Network/Campus Operations Center and serve as the eyes and voice of the site. This role is responsible for continuous 24/7 monitoring of xAI data center campuses, detecting and verifying campus-impacting events, assembling the right responders, and running incident communications that leadership can trust.
Key responsibilities include:
**Continuous Monitoring**: Staff the console per shift schedule to maintain 24/7 coverage (2 technicians per site). Monitor cluster health dashboards, node availability, network health, facility trend panels (power/cooling), storage alarms, and threshold breaches. Acknowledge every page/alert within SLA, classify it (actionable/known/noise), and log disposition. Maintain awareness of ongoing maintenance and planned work.
**Detection, Triage & Escalation**: Detect, verify, and escalate within defined time budgets. Operate the escalation matrix (NOC → on-call SRE → domain owners). Recommend incident declaration and severity; declare directly per runbook when thresholds are clear.
**Incident Communications & Coordination**: Open and run incident bridges; get the right people engaged within SLA. Own stakeholder communications with first update and fixed cadence until resolution. Maintain real-time incident timeline with timestamps, actions, decisions, and engagements. Track ownership and call out stalls.
**RCA Framing & Closure**: Produce initial framing for major outages (what happened, when, impact, recent changes, engaged teams). Hand framing to SRE/Hardware FA for deep analysis. Write major-incident reports and open corrective projects in Linear with named owners, tracking to closure.
**Shift Operations & Improvement**: Run structured shift handoffs and maintain durable shift logs. Own and continuously improve NOC runbooks (escalation matrix, comms templates, severity ladders, response procedures). Participate in game days; every incident where the runbook was wrong produces a change before closure.
This role explicitly does NOT include wrench work, power/cooling/building operations, deep hardware analysis, monitoring architecture design, or reliability tooling development—those belong to specialized teams.
The ideal candidate thrives under pressure, communicates clearly in real time, recognizes patterns across domains (compute, network, storage, power/cooling), and maintains institutional memory across shifts and sites.