SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Graphcore, part of the SoftBank Group, is a leading innovator in AI compute hardware and infrastructure. We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for our Arm-based hardware platforms across lab and data center environments.
You will focus on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. Working closely with engineering, platform, and data center teams, you'll ensure the reliability and performance of next-generation AI systems.
Key responsibilities include:
- Lead advanced break-fix troubleshooting for server blades, motherboards, power systems, and rack-scale infrastructure
- Support engineering bring-up activities, including component validation and firmware interaction testing
- Diagnose system-level failures involving thermal behavior, power anomalies, network configuration, and BIOS/BMC issues
- Collaborate with server engineering teams on root cause analysis and design improvements
- Support deployment and rollout of next-generation hardware platforms through validation and qualification cycles
- Interface with facilities and infrastructure teams to understand environmental factors impacting system reliability
- Develop and maintain standard operating procedures, troubleshooting guides, and validation documentation
- Provide guidance and mentorship to junior technicians and engineers on troubleshooting methodologies
- Participate in on-call rotations during critical engineering milestones or hardware bring-up phases
Required qualifications:
- Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or related discipline
- 10+ years experience with server hardware architectures and board-level debugging
- Experience analyzing system logs, hardware telemetry, and power/thermal metrics to isolate hardware failures
- Hands-on experience with HPC systems, AI compute platforms, or rack-scale infrastructure
- Strong collaboration skills and ability to work effectively in fast-paced engineering environments
- Excellent written and verbal communication skills
Desirable experience includes prototype/pre-production hardware bring-up, data center facilities knowledge (liquid cooling, power distribution), Python/Bash automation for hardware validation, and structured failure analysis methodologies.