SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Graphcore, a SoftBank-backed leader in AI compute infrastructure, is seeking a Senior Systems Engineer to support the bring-up, validation, and operational reliability of its Arm-based hardware platforms across lab and data center environments.
You will provide advanced operational, diagnostic, and engineering support for complex AI compute systems, including server blades, racks, and rack-scale infrastructure. This role bridges hardware engineering, platform teams, and data center operations to ensure next-generation AI systems perform reliably from early development through production deployment.
Key responsibilities include:
- Lead advanced troubleshooting for server blades, motherboards, power systems, and rack-scale infrastructure
- Support hardware bring-up activities, component validation, and firmware interaction testing
- Diagnose system-level failures involving thermal behavior, power anomalies, network configuration, and BIOS/BMC issues
- Collaborate with server engineering teams on root cause analysis and design improvements
- Support deployment and qualification cycles for next-generation hardware platforms
- Interface with facilities and infrastructure teams on environmental factors affecting reliability
- Develop and maintain standard operating procedures, troubleshooting guides, and validation documentation
- Mentor junior technicians and engineers on troubleshooting methodologies and hardware diagnostics
- Participate in on-call rotations during critical engineering milestones
Required: Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or related field; 7+ years with server hardware architectures and board-level debugging; experience analyzing system logs, hardware telemetry, and power/thermal metrics; hands-on experience with HPC systems, AI compute platforms, or rack-scale infrastructure; strong collaboration and communication skills.
Desirable: prototype/pre-production hardware bring-up experience; familiarity with data center facilities (liquid cooling, power distribution); Python, Bash, or automation tools for hardware validation; structured failure analysis and reliability engineering methodologies.