SlipstreamJobsFresh Startup & VC-Backed Jobs

Data Center Hardware Quality & Reliability Engineer

OpenAI - San Francisco, CA, USA - In-office - posted 2026-09-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

OpenAI is seeking a Data Center Hardware Quality & Reliability Engineer to own the end-to-end hardware quality and reliability lifecycle for its data-center infrastructure, spanning both third-party and first-party platforms. This is the first hire in this function and requires someone who can combine practical hardware/system understanding, reliability engineering, data fluency, and cross-functional technical leadership at data-center scale. Key responsibilities include: • Build and govern the field-quality data model integrating telemetry, tickets, RMA/repair, failure analysis, firmware, configuration, supplier, and manufacturing genealogy data. • Define and track critical reliability metrics (AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, forecast-versus-actual) with explicit denominators and uncertainty quantification. • Provide both fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations. • Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, corrective-action verification, and recurrence monitoring. • Develop cohort, life-data, reliability-growth, and spare-demand projections segmented by product, FRU, supplier, configuration, geography, and age. • Partner with Manufacturing Quality Engineering (MQE) and New Product Introduction (NPI) teams to convert field failure mechanisms into manufacturing-test coverage, screening/stress profiles, diagnostics, control plans, DFR/DFS requirements, FMEA/FTA, mission profiles, FRU strategy, and qualification gates. • Verify whether upstream changes reduce field recurrence. • Define supplier and contract manufacturer (CM) failure analysis standards, field-data contracts, scorecards, escalation paths, and closure evidence requirements. • Provide serviceability, total cost of ownership (TCO), and spares inputs without owning inventory execution or procurement. • Create executive decision packages summarizing population at risk, exposure, confidence, options, cost/risk trade-offs, and recommendations. • Run the cross-functional reliability council and mentor field quality engineers as the team grows. Requirements: • BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience; MS preferred. • 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure; 3+ years owning field-failure, RMA, or CAPA outcomes. • Solid working understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, manufacturing test, and fleet behavior; deep expertise in every subsystem is not required. • Reliability statistics: censored life data, Weibull/Poisson/binomial methods, confidence bounds, MTBF/MTTR, and reliability growth modeling. • Hands-on experience with FMEA/FTA, accelerated or reliability-demonstration testing, 8D/CAPA, failure analysis, and corrective-action verification. • Working proficiency with SQL and Python/R or equivalent analytics tools. • Ability to influence design, validation, operations, suppliers/CMs, and senior leaders without direct authority. Preferred skills: GPU/AI server platforms, liquid cooling, high-power delivery, high-speed networking, rack integration, or data-center operations; design for serviceability (FRU boundaries, diagnostics, repair workflows, tooling/access, spares policy); qualification-to-field correlation and mission-profile development; ODM/CM/supplier experience (FA quality, audit, QBR, corrective-action governance); Linux/BMC/IPMI/Redfish logs and fleet telemetry; leadership of cross-generation reliability programs or launch-readiness gates.

Similar roles