SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Graphcore, part of the SoftBank Group, is seeking a Staff System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. This role is distinct from traditional validation positions—rather than executing test plans, you will develop the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability during bring-up and validation.
You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Your responsibilities include:
• Learning and extending existing Arm diagnostics tools to support Graphcore's AI platform
• Developing specialized diagnostics and stress-testing software for CPU, memory, interconnect, PCIe, and other system components
• Building reusable validation utilities that integrate into the existing validation framework
• Developing tools that expose intermittent hardware failures, including silent data corruption (SDC), timing-related issues, and hardware instability
• Designing configurable diagnostic applications supporting multiple silicon revisions, platforms, and execution environments
• Developing automation for executing diagnostics across server and rack-scale validation environments
• Collaborating with hardware architects to understand new hardware capabilities and identify missing diagnostics
• Working closely with firmware and validation teams to improve hardware observability
• Debugging low-level hardware, firmware, operating system, and driver interactions
• Analyzing diagnostic output and improving fault isolation methodologies
• Developing software that enables engineers to reproduce and investigate difficult hardware failures
• Contributing reusable libraries and utilities that improve engineering productivity across multiple validation teams
Success in this role means extending existing diagnostics technologies to support new hardware, developing new tools for emerging capabilities, building reusable software components that integrate seamlessly with the validation framework, improving hardware observability and root-cause isolation, detecting intermittent failures that traditional techniques cannot expose, and delivering configurable diagnostics that scale across multiple platforms and silicon revisions.
This position offers a unique opportunity to work at the intersection of hardware architecture, diagnostics, and software engineering while helping define how next-generation AI systems are validated. Rather than maintaining an existing framework, you will develop the specialized diagnostics and stress tools that make that framework valuable, directly influencing hardware quality, silicon bring-up, and long-term platform reliability.
REQUIREMENTS:
Essential:
• Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or related technical discipline
• 10+ years of experience developing hardware diagnostics, system validation software, firmware validation tools, or low-level systems software
• Strong software development skills in Python and C/C++
• Strong Linux systems experience
• Strong understanding of modern server platform architecture (CPU, memory, PCIe, storage, networking, firmware, OS, device drivers)
• Experience debugging hardware/software interactions across multiple platform layers
• Experience developing diagnostics, stress testing, or hardware validation software
• Experience with server platforms, embedded systems, or SoC validation
• Strong analytical and debugging skills
• Experience collaborating across hardware, firmware, software, and validation teams
• Excellent communication and problem-solving skills
Desirable:
• Arm-based server platforms
• AI accelerators or high-performance computing systems
• Silent Data Corruption (SDC) testing
• Hardware stress testing
• Power transient analysis
• Hardware fault injection methodologies
• Firmware validation
• Silicon bring-up
• Performance characterization
• Hardware telemetry and observability
• BMC firmware or server management
• RAS (Reliability, Availability, Serviceability) technologies
• Validation across subsystems: CPU scaling/cache behavior, memory (DDR/HBM) bandwidth/latency/NUMA effects, interconnect contention, PCIe/I-O throughput/latency, high-speed I/O validation