SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Graphcore is building the future of AI compute infrastructure. As part of the SoftBank Group, the company develops the complete AI compute stack—from silicon and software to datacenter-scale infrastructure. This role sits within the Performance & Reliability team, which is responsible for determining whether entire racks of machines meet production standards.
You will measure and evaluate large-scale Linux systems, from single racks to datacenter scale. The work is not limited to running benchmarks; you'll determine whether systems behave correctly and are reliable enough for production deployment. Your responsibilities span designing workloads, building execution systems, and interpreting results to drive engineering decisions.
The role offers breadth across infrastructure, workloads, and analysis. While you may specialize in one area, the team collectively owns understanding system behavior at scale with no gaps. You'll expand measurement coverage from small clusters to full racks, design workloads that expose system behavior, build systems to run experiments across large clusters, and interpret results to define production readiness criteria.
This is a hybrid systems engineering, measurement, and judgment role. You'll work in areas where requirements aren't fully defined and careful measurement matters more than output volume.
Essential qualifications: strong software engineering experience across multiple projects over several years; hands-on experience in Linux-based environments, ideally with distributed or high-performance systems; Python proficiency; experience with automation and CI/CD systems (GitLab CI, Jenkins, GitHub Actions); ability to design, implement, and run experiments producing meaningful results; skill in interpreting results and communicating findings clearly for decision-making; comfort working in ambiguous areas requiring engineering judgment.
Desirable: experience with large-scale or distributed systems (clusters, cloud platforms, HPC); performance, reliability, or systems-level testing/measurement background; familiarity with pytest or similar test frameworks; experience analyzing system behavior under load (compute, network, or ML workloads); exposure to containerization or orchestration tools.