SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Graphcore, a SoftBank-backed AI infrastructure innovator, is seeking a Staff Hardware Engineer to lead advanced operational and diagnostic support for its Arm-based AI compute platforms. This role sits at the intersection of hardware engineering and systems operations, focusing on the bring-up, validation, and troubleshooting of complex AI compute systems including server blades, racks, and rack-scale infrastructure deployed across lab and data center environments.
You will own advanced break-fix troubleshooting for server hardware, motherboards, power systems, and rack-scale infrastructure. Key responsibilities include supporting engineering bring-up activities with component validation and firmware interaction testing; diagnosing system-level failures involving thermal behavior, power anomalies, network configuration, and BIOS/BMC issues; and collaborating with server engineering teams on root cause analysis and design improvements.
You will support deployment and rollout of next-generation hardware platforms through structured validation and qualification cycles, interface with facilities and infrastructure teams on environmental factors affecting reliability, and develop standard operating procedures, troubleshooting guides, and validation documentation. A critical aspect of this role is providing guidance and mentorship to junior technicians and engineers on troubleshooting methodologies and hardware diagnostics, plus participation in on-call rotations during critical engineering milestones.
The ideal candidate brings 8+ years of hands-on experience with server hardware architectures and board-level debugging, with demonstrated expertise analyzing system logs, hardware telemetry, and power/thermal metrics. Experience with HPC systems, AI compute platforms, or rack-scale infrastructure is essential. Desirable qualifications include prototype/pre-production hardware bring-up experience, familiarity with data center facilities (liquid cooling, power distribution), scripting skills (Python, Bash), and exposure to structured failure analysis and reliability engineering methodologies.