SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
xAI is seeking a Hardware Deployment Engineer Lead to own end-to-end bring-up of GPU compute hardware across the world's largest AI training clusters. You will build and lead a dedicated in-house hardware deployment team responsible for L11 integration, hardware bring-up, and post-L11 repair of GB300-class systems across multiple data halls concurrently. Your team's throughput directly determines how fast xAI can compute online—this is one of the most critical path activities in the company.
Key responsibilities include leading, hiring, and developing a dedicated hardware deployment team with full ownership of team structure and staffing. You will own L11 rack integration and compute hardware bring-up across multiple data halls concurrently, from delivery dock to healthy production handoff. You'll drive aggressive bring-up timelines to achieve 95%+ node availability within days of rack delivery and 100% closure within one week per data hall.
You will own post-L11 hardware health by running systematic health pushes to sustain greater than 98% node availability prior to turnover to operations. You'll internalize non-RMA hardware repairs to maximize hardware recovery, minimize repair backlogs, and reduce dependence on OEM turnaround times. You will develop and enforce vendor SLAs for OEM and supplier responsibilities, preventing accumulation of unrepaired hardware and repair backlogs.
Additional responsibilities include performing root cause analysis of hardware failures discovered during L11 and driving corrective actions with vendors and internal engineering teams. You'll partner with site operations on hardware debugging and repair, and train site operations teams to support future data center deployments. You will build, document, and continuously improve deployment processes, tooling, and training so bring-up capability scales across sites and future hardware generations.
Required qualifications: 5+ years of hands-on experience deploying, integrating, or repairing compute/server hardware at data center scale; direct experience with L11 (rack-level) integration and bring-up of GPU or accelerator-based systems; demonstrated experience leading technician or engineering teams in a fast-paced deployment, manufacturing, or data center environment; deep troubleshooting skills across servers, GPUs, NVLink/fabric interconnects, high-speed networking, and liquid cooling systems; willingness to work on-site in Memphis, TN, including extended hours and weekends during critical bring-up phases.
Preferred experience includes NVIDIA GB200/GB300 NVL72 or similar rack-scale liquid-cooled GPU systems, standing up a new team or function from scratch, managing OEM/ODM vendor relationships, hardware failure analysis and RMA processes, data center automation and hardware health telemetry, and a track record of driving step-change improvements in deployment velocity or cost.