SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Graphcore is seeking a Staff Hardware Engineer to own the reliability and operational excellence of AI compute hardware as it scales from lab discovery to data centre deployment. You will lead advanced diagnostic, validation, and engineering support across lab and production environments, working at the intersection of hardware bring-up, systems integration, and operational resilience.
In this hands-on role, you will diagnose complex failures spanning server blades, racks, power systems, thermal behaviour, network configuration, and firmware/BMC issues. You will conduct structured root cause analysis, develop corrective actions, and establish better operating practices that improve hardware reliability and accelerate platform deployment. Your work directly impacts how quickly Graphcore can move new AI compute platforms from early validation through confident production deployment.
You will collaborate closely with Systems Engineering, Hardware Engineering, server engineering, firmware teams, platform architects, and data centre operations. The team values speed, direct ownership, data-driven decision-making, and clear communication. You will guide junior engineers, improve diagnostic documentation, and raise the standard of hardware validation across the organization.
Required experience includes strong knowledge of server hardware architectures and board-level debugging, hands-on isolation of failures using system logs, telemetry, power data, and thermal metrics, and experience with HPC systems, AI compute platforms, or rack-scale infrastructure. You should be comfortable leading structured problem-solving across engineering and operations teams, with clear written and verbal communication skills and the judgement to mentor others.
Graphcore, part of the SoftBank Group, is building the hardware and systems infrastructure for next-generation AI breakthroughs. The company brings together AI research specialists, silicon designers, software engineers, and systems architects to solve complex problems in AI compute infrastructure.