SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
TensorWave is building a cloud platform for AI compute infrastructure. The Hardware Diagnostics Engineer role owns the validation, troubleshooting, and lifecycle management of GPU servers before they reach production.
Key Responsibilities:
- Run burn-in and stress testing on servers and GPUs, interpret results, and determine production readiness
- Triage hardware failures across GPUs, memory, drives, NICs, PSUs, and cabling; reproduce issues, isolate faulty components, and document findings
- Manage out-of-band server access via IPMI and Redfish for power control, boot configuration, BIOS settings, and log collection
- Apply firmware updates across the fleet following qualified baselines and rollout procedures
- Drive RMAs with vendors from ticket creation through replacement, installation, and return of failed parts
- Maintain accurate asset, serial, and replacement history in NetBox
- Identify and escalate failure patterns across the fleet
- Improve runbooks and script repetitive troubleshooting steps
- Partner with datacenter operations during turn-ups and expansions
- Participate in on-call rotation for hardware escalations
The role reports into infrastructure engineering and is hands-on: you'll spend time in the datacenter, at the rack, and working through vendor relationships. By 90 days, you'll run burn-in cycles and drive RMAs independently. By six months, you'll spot failure patterns early and have automated at least one manual process.
Requirements:
- 3–6 years in datacenter operations, systems administration, hardware support, or infrastructure engineering
- Hands-on experience with enterprise server hardware: component replacement, POST and boot failures, rack-level troubleshooting
- Practical experience with BMCs and out-of-band management (IPMI, Redfish, iDRAC, iLO, or equivalent)
- Strong Linux troubleshooting: boot process, driver/device issues, diagnostic tools (dmesg, lspci, ipmitool, SMART)
- Ability to read sensor data, event logs, thermal and power telemetry to distinguish real failures from noise
- Working scripting ability in Bash or Python to automate repetitive tasks and read existing tooling
- Experience running hardware RMAs with vendors or proven track record of driving issues to closure with external parties
- Methodical troubleshooting approach: isolate variables, document evidence, avoid multi-variable changes
- Clear written communication for tickets, runbooks, and vendor cases
Preferred Qualifications:
- GPU server experience, especially AMD GPUs and ROCm
- Burn-in, stress testing, or node validation tooling in GPU or HPC environments
- Familiarity with firmware update processes and fleet-wide staging
- NetBox or other DCIM/IPAM tooling
- Ansible or Python against REST APIs
- Prior work in high-volume hardware environments (hyperscaler, colo, integrator, manufacturing test)