SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Graphcore, part of the SoftBank Group, is a leading innovator in AI compute hardware and infrastructure. This Staff Hardware Engineer role supports the operational reliability and bring-up of Graphcore's next-generation Arm-based AI compute platforms across lab and data center environments.
You will lead advanced troubleshooting and validation of complex hardware systems, including server blades, motherboards, power systems, and rack-scale infrastructure. Key responsibilities include:
• Lead break-fix troubleshooting for server hardware, power systems, and rack-scale infrastructure
• Support engineering bring-up activities, including component validation and firmware interaction testing
• Diagnose system-level failures involving thermal behavior, power anomalies, network configuration, and BIOS/BMC issues
• Collaborate with server engineering teams on root cause analysis and design improvements
• Support deployment and rollout of next-generation hardware through structured validation cycles
• Interface with facilities and infrastructure teams on environmental factors affecting system reliability
• Develop and maintain standard operating procedures, troubleshooting guides, and validation documentation
• Provide guidance and mentorship to junior technicians and engineers on troubleshooting methodologies
• Participate in on-call rotations during critical engineering milestones and hardware bring-up phases
Required: Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or related field; 12+ years of experience with server hardware architectures and board-level debugging; experience analyzing system logs, hardware telemetry, and power/thermal metrics; hands-on experience with HPC systems, AI compute platforms, or rack-scale infrastructure; strong collaboration and communication skills.
Desirable: prototype or pre-production hardware bring-up experience; familiarity with data center facilities including liquid cooling and power distribution; Python, Bash, or automation tools for hardware validation; structured failure analysis and reliability engineering methodologies.