SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Hardware Engineer

Graphcore - Austin, TX, United States - In-office - posted 2026-09-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Graphcore, part of the SoftBank Group, is a leading innovator in AI compute hardware and infrastructure. The company develops hardware, software, and systems infrastructure to unlock next-generation AI breakthroughs and power widespread AI adoption across industries. We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore's Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. You will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore's AI infrastructure platforms, working closely with server engineering, firmware teams, platform architects, and data center operations. Key Responsibilities: - Lead advanced break-fix troubleshooting for server blades, motherboards, power systems, and rack-scale infrastructure - Support engineering bring-up activities, including component validation and firmware interaction testing - Diagnose system-level failures involving thermal behavior, power anomalies, network configuration, and BIOS/BMC issues - Collaborate with server engineering teams to perform root cause analysis and propose corrective actions or design improvements - Support deployment and rollout of next-generation hardware platforms through structured validation and qualification cycles - Interface with facilities and infrastructure teams to understand environmental factors impacting system reliability - Develop and maintain standard operating procedures (SOPs), troubleshooting guides, and validation documentation - Provide guidance and mentorship to junior technicians and engineers on troubleshooting methodologies and hardware diagnostics - Participate in on-call rotations or off-hours support during critical engineering milestones or hardware bring-up phases Requirements: Essential: - Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or related discipline - 10 years of experience with server hardware architectures and board-level debugging - Experience analyzing system logs, hardware telemetry, and power/thermal metrics to isolate hardware failures - Hands-on experience with HPC systems, AI compute platforms, or rack-scale infrastructure - Strong collaboration skills and ability to work effectively in fast-paced engineering environments - Excellent written and verbal communication skills Desirable: - Experience supporting prototype or pre-production hardware bring-up - Familiarity with data center facilities, including liquid cooling and power distribution systems - Experience using Python, Bash, or automation tools for hardware validation or troubleshooting - Exposure to structured failure analysis and reliability engineering methodologies

Similar roles