SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Graphcore, a SoftBank Group company and leader in AI compute infrastructure, is seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for its Arm-based hardware platforms. This role focuses on hardware bring-up, validation, and troubleshooting of complex AI compute systems, including server blades, racks, and rack-scale infrastructure deployed in lab and data center environments.
You will lead advanced break-fix troubleshooting for server hardware, support engineering bring-up activities including component validation and firmware interaction testing, and diagnose system-level failures involving thermal behavior, power anomalies, network configuration, and BIOS/BMC issues. Working closely with server engineering, platform, and data center operations teams, you'll perform root cause analysis, propose design improvements, and support deployment and rollout of next-generation hardware platforms through structured validation cycles.
Key responsibilities include developing and maintaining standard operating procedures, troubleshooting guides, and validation documentation; providing guidance and mentorship to junior technicians and engineers on troubleshooting methodologies; interfacing with facilities teams to understand environmental factors impacting reliability; and participating in on-call rotations during critical engineering milestones.
Required qualifications: Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or related field; 7+ years of experience with server hardware architectures and board-level debugging; experience analyzing system logs, hardware telemetry, and power/thermal metrics; hands-on experience with HPC systems, AI compute platforms, or rack-scale infrastructure; and strong collaboration and communication skills.
Desirable experience includes prototype or pre-production hardware bring-up, familiarity with data center facilities (liquid cooling, power distribution), scripting with Python or Bash for hardware validation, and exposure to structured failure analysis and reliability engineering methodologies.