SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Graphcore, a SoftBank Group company and leader in AI compute infrastructure, is seeking a Senior Principal Network Engineer to design, deploy, and optimize next-generation AI data center networks. This role partners closely with the Network Architecture Lead to design and scale high-performance computing (HPC) network fabrics supporting GPU clusters, working across hardware, networking, and AI application layers.
Key responsibilities include defining ultra-high-bandwidth, non-blocking AI network fabrics using Clos spine-leaf-super-spine architectures for large-scale distributed AI workloads. You will optimize performance of lossless Ethernet fabrics using congestion control mechanisms such as PFC, ECN, and DCQCN to support RDMA/RoCEv2 communication. The role involves leading NetDevOps initiatives to implement automation for provisioning, configuration management, and network remediation across data center infrastructure.
You will design and deploy high-resolution telemetry pipelines to monitor network health, detect microbursts, and analyze congestion patterns. This includes supporting modeling, deployment, configuration, and monitoring of data center network fabrics including scale-out, scale-up, and front-end networks. Cross-functional collaboration with hardware engineers, AI researchers, and data center operations teams is essential to co-design high-performance infrastructure.
The role includes providing technical leadership and mentorship to network engineers while establishing best practices and operational standards. You will contribute to long-term networking strategy and roadmap for Graphcore's AI infrastructure and research next-generation high-speed networking technologies and vendor solutions.
Required qualifications: 12+ years of progressive network engineering experience with at least 3 years in hyperscale, high-density, or HPC data center environments. Expert-level knowledge of data center routing and switching protocols (BGP, OSPF, EVPN-VXLAN), strong operational understanding of RDMA technologies (RoCEv2, InfiniBand), and hands-on experience with modern merchant silicon platforms (Arista EOS, Cisco NX-OS, SONiC). Experience deploying 400G/800G optics and large-scale fabric architectures required. Proficiency in Python, Go, Bash, or similar automation languages essential.