SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer - Datacenter

xAI - Memphis, TN, United States - In-office - posted 2026-08-13

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

xAI is seeking a Site Reliability Engineer focused on Hardware to support its datacenter operations in Memphis, Tennessee. This role combines hardware reliability engineering with vendor management and firmware analysis to ensure the stability and performance of xAI's AI infrastructure. You will serve as the technical expert responsible for firmware evaluation, hardware specifications review, and failure diagnosis. Key responsibilities include analyzing firmware packages and hardware specifications for compatibility, performance, and reliability before deployment; investigating and diagnosing complex hardware failures, including intermittent or ambiguous issues ("grey failures"); managing vendor relationships and RMA (Return Merchandise Authorization) processes; collaborating with datacenter operations technicians on real-time troubleshooting; developing monitoring tools and scripts to detect hardware anomalies early; documenting failure modes, root cause analyses, and reliability models; and participating in on-call rotations for hardware incidents. You will need a bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or equivalent experience, plus 2+ years in hardware reliability engineering, ideally in high-performance computing or datacenter environments. Required expertise includes firmware analysis, hardware specifications review, RMA process management, and the ability to diagnose complex hardware failures using diagnostic tools and logic analyzers. You should be familiar with datacenter hardware (servers, GPUs, networking equipment) and proficient in scripting (Python, Bash) plus at least one systems language (C, C++, Java, Rust, or similar). Preferred qualifications include experience in AI/ML infrastructure or supercomputing, knowledge of vendor ecosystems (NVIDIA, Dell, HP, Supermicro), hardware engineering certifications (CRE, CompTIA Server+), and prior work in fast-paced startups or tech companies. The organization operates with a flat structure, emphasizing hands-on contribution, engineering excellence, strong communication, and prioritization skills.

Similar roles