SlipstreamJobsFresh Startup & VC-Backed Jobs

Hardware Failure Analysis Engineer

xAI - Memphis, TN, United States - In-office - posted 2026-09-03

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

xAI is seeking a Hardware Failure Analysis Engineer to support its data center operations in Memphis, Tennessee. In this role, you will serve as a technical expert in firmware analysis, hardware specifications, vendor relations, and failure diagnosis for high-performance computing infrastructure. Key responsibilities include analyzing firmware packages and hardware specifications for compatibility, performance, and reliability before deployment. You will investigate and diagnose complex hardware failures, including intermittent or ambiguous "grey failures," using rigorous testing and data-driven analysis. You'll manage vendor relationships and RMA (Return Merchandise Authorization) processes, negotiating resolutions and holding vendors accountable. You will collaborate closely with Data Center Operations Technicians to troubleshoot and optimize hardware systems in real-time. The role requires developing monitoring tools, scripts, and processes to detect hardware anomalies early and minimize downtime. You'll document failure modes, root cause analyses, reliability models, and RMA outcomes into a team knowledge base, and participate in on-call rotations for incident response. Required qualifications include a Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or equivalent experience, plus 2+ years in hardware reliability engineering, preferably in high-performance computing or data center environments. You must have proven expertise in firmware analysis, hardware specifications review, and release validation. Strong experience with RMA processes, vendor negotiations, and diagnosing complex hardware failures is essential. Familiarity with data center hardware (servers, GPUs, networking equipment) and proficiency in scripting (Python, Bash) plus at least one systems language (C, C++, Java, Rust) are required. Preferred experience includes work in AI/ML infrastructure or supercomputing environments, knowledge of vendor ecosystems (NVIDIA, Dell, HP, Supermicro), hardware engineering certifications (CRE, CompTIA Server+), and prior startup experience.

Similar roles