SlipstreamJobsFresh Startup & VC-Backed Jobs

Director, Site Operations

xAI - Memphis, TN, United States - In-office - posted 2026-09-15

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

xAI is seeking a Director of Site Operations to own cluster uptime and reliability for its advanced AI supercompute infrastructure. This is an extreme ownership role managing 250+ personnel across 5+ data center sites operating 24/7, including site managers, shift supervisors, technicians, and a site reliability engineering team. Key responsibilities include: - Serving as the extreme owner of node, rack, and cluster health across all sites, accountable for customer Service Level Agreements and exceptional uptime - Leading a large, multi-site operations organization across four 24/7 shifts, building a culture of excellence and accountability - Driving systematic node and rack remediation, minimizing mean time to repair through command-line and physical intervention - Coordinating with facilities, network engineering, and tenant representatives to limit downtime from power, cooling, and maintenance issues - Directing vendor execution for hardware rework and field operations to ensure repairs and capacity work happen at required speed - Owning the SRE organization responsible for proactive cluster health monitoring, reactive fault mitigation, root cause analysis, and reliability procedures - Leading data-driven improvement initiatives using operational metrics to optimize team resources and uptime across sites - Setting the standard for incident response during cluster-impacting events with clear direction and tight communication - Scaling operations and standardizing best practices as the company's footprint expands The ideal candidate will thrive in a dynamic, mission-focused environment with hands-on leadership style. The role requires frequent travel to data center locations and willingness to work extended hours and weekends as needed. Physical capability to handle data center tasks (lifting up to 50 lbs, standing for extended periods, occasional ladder use) is required. QUALIFICATIONS: - Bachelor's degree and 7+ years of large-scale operations experience with 5+ years leading people leaders of technical teams, OR 10+ years of large-scale operations experience with 5+ years leading people leaders of technical teams - Proven ability to lead large, multi-site, 24/7 operations in fast-paced, high-responsibility settings - Deep expertise in server hardware, cluster reliability, and data center technologies from deployment through lifecycle management - Experience supporting compute-heavy environments (AI, machine learning, high-performance computing) at scale - Track record of owning uptime, SLAs, or reliability metrics for large compute clusters - Experience leading site reliability engineering or equivalent reliability-focused teams with root cause analysis and procedure ownership - Strong analytical skills and ability to explain technical concepts to diverse audiences - History of partnering with vendors at scale, driving MTTR improvements, and scaling operations across multiple sites - Familiarity with monitoring and automation tooling (Jira, Python, Bash) - Ability to thrive in dynamic, mission-focused environment with on-call ownership of cluster-impacting events

Similar roles