SlipstreamJobsFresh Startup & VC-Backed Jobs

Network Operations Center Specialist

xAI - Memphis, TN, United States - In-office - posted 2026-09-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

xAI is seeking a Network Operations Center (NOC) Specialist to serve as the eyes and voice of the campus, monitoring infrastructure health around the clock and orchestrating incident response. This is a communications-and-judgment role, not a junior engineering position. You will staff the NOC console per shift schedule, watching cluster health, node availability, network health, facility trends, storage alarms, and threshold breaches. You acknowledge every page within SLA, classify alerts (actionable/known/noise), and log disposition. You detect, verify, and escalate incidents within time budgets, operating the escalation matrix to page on-call SRE and domain owners correctly the first time. You open and run incident bridges, owning stakeholder communications with first updates within SLA and fixed-cadence follow-ups. You maintain incident timelines in real time and call out ownership stalls. You produce first-pass RCA framing (what happened, when, impact, who's engaged) and hand it to SRE and Hardware Failure Analysis for deeper analysis—the NOC does not publish root cause. You run structured shift handoffs, maintain durable shift logs, and preserve cross-site awareness. You write major-incident reports, open corrective projects in Linear, and chase them to closure. You maintain and continuously improve NOC runbooks, escalation matrices, and communications templates, and participate in SRE-run game days. This role explicitly does not include wrench work, plant operation, deep root-cause analysis, monitoring design, technical SEV command, or tool building—those belong to SiteOps, Facilities, Hardware Failure Analysis, Site SRE, and Software Platforms. Required: 24/7 operations experience (NOC, SOC, dispatch, mission control, or equivalent); proven ability to acknowledge, classify, and escalate incidents under SLA in high-signal environments; experience opening and running incident bridges with stakeholder updates and live timeline hygiene; excellent written and verbal communication; pattern recognition across compute, network, storage, and/or facilities signals; experience following and improving operational processes; willingness to work rotating shifts including nights and weekends. Preferred: prior NOC, data center operations, or campus reliability experience in high-performance computing, AI/ML infrastructure, or large-scale production; experience writing major-incident reports and driving corrective follow-ups; familiarity with Linear or similar work-tracking tools; experience partnering with SRE, SiteOps, and Facilities; participation in game days or runbook improvement programs; prior work in fast-paced startups or tech companies.

Similar roles