SlipstreamJobsFresh Startup & VC-Backed Jobs

Hardware Diagnostics Engineer - Infrastructure

TensorWave - Las Vegas, NV, United States - In-office - posted 2026-10-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

TensorWave is building a cloud platform for AI compute infrastructure. The Hardware Diagnostics Engineer role owns the validation, troubleshooting, and lifecycle management of GPU servers before they reach production. Key Responsibilities: - Run burn-in and stress testing on servers and GPUs, interpret results, and determine production readiness - Triage hardware failures across GPUs, memory, drives, NICs, PSUs, and cabling; reproduce issues, isolate faulty components, and document findings - Manage out-of-band server access via IPMI and Redfish for power control, boot configuration, BIOS settings, and log collection - Apply firmware updates across the fleet following qualified baselines and rollout procedures - Drive RMAs with vendors from ticket creation through replacement, installation, and return of failed parts - Maintain accurate asset, serial, and replacement history in NetBox - Identify and escalate failure patterns across the fleet - Improve runbooks and script repetitive troubleshooting steps - Partner with datacenter operations during turn-ups and expansions - Participate in on-call rotation for hardware escalations The role reports into infrastructure engineering and is hands-on: you'll spend time in the datacenter, at the rack, and working through vendor relationships. By 90 days, you'll run burn-in cycles and drive RMAs independently. By six months, you'll spot failure patterns early and have automated at least one manual process. Requirements: - 3–6 years in datacenter operations, systems administration, hardware support, or infrastructure engineering - Hands-on experience with enterprise server hardware: component replacement, POST and boot failures, rack-level troubleshooting - Practical experience with BMCs and out-of-band management (IPMI, Redfish, iDRAC, iLO, or equivalent) - Strong Linux troubleshooting: boot process, driver/device issues, diagnostic tools (dmesg, lspci, ipmitool, SMART) - Ability to read sensor data, event logs, thermal and power telemetry to distinguish real failures from noise - Working scripting ability in Bash or Python to automate repetitive tasks and read existing tooling - Experience running hardware RMAs with vendors or proven track record of driving issues to closure with external parties - Methodical troubleshooting approach: isolate variables, document evidence, avoid multi-variable changes - Clear written communication for tickets, runbooks, and vendor cases Preferred Qualifications: - GPU server experience, especially AMD GPUs and ROCm - Burn-in, stress testing, or node validation tooling in GPU or HPC environments - Familiarity with firmware update processes and fleet-wide staging - NetBox or other DCIM/IPAM tooling - Ansible or Python against REST APIs - Prior work in high-volume hardware environments (hyperscaler, colo, integrator, manufacturing test)

Similar roles