SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Graphcore, a leading AI compute innovator and member of the SoftBank Group, is seeking a Principal Reliability Scientist to lead reliability activities across complex, high-performance AI hardware systems. Reporting to Quality leadership within Manufacturing Operations, you will work with established reliability experts and cross-functional teams to use experimental data and advanced modelling to inform design decisions, validate product reliability, and optimize serviceability strategies.
Key responsibilities include defining and refining reliability requirements across silicon, board, and system levels in partnership with research and design teams. You will apply advanced reliability methodologies to highly innovative systems, including challenges associated with liquid-cooled architectures and fluid dynamics. The role involves designing and executing experiments to generate high-quality reliability and performance data with statistical rigor, analyzing experimental, field, and manufacturing data to quantify reliability metrics such as MTBF, MTTR, RAS characteristics, and soft error rates (SER).
You will use data-driven insights to inform product design trade-offs, reliability targets, and spares provisioning strategies. Collaboration with chip, board, and system design teams is essential to influence architecture and component selection based on reliability considerations. You will support development of system-level reliability models incorporating thermal, mechanical, and fluid behavior, lead complex root cause investigations into reliability issues, and drive corrective and preventative actions across teams.
Additionally, you will contribute to the evolution of reliability tools, processes, and best practices within the organization, and communicate complex reliability concepts, risks, and recommendations clearly to a wide range of stakeholders.
Required qualifications include a strong background in reliability engineering or reliability science within semiconductor, hardware, or complex systems environments; experience with physics-of-failure approaches in high-performance computing or AI hardware; expertise in reliability modelling, experimental design, and statistical data analysis; proven ability to work with and interpret experimental reliability data; knowledge of key reliability metrics; ability to operate effectively in complex, cross-functional environments; strong problem-solving skills; and excellent communication abilities.
Preferred qualifications include experience with liquid cooling systems, fluid dynamics, or thermally complex hardware environments; knowledge of soft error mechanisms and SER modeling; and experience contributing to reliability strategy, processes, or tooling improvements.