SlipstreamJobsFresh Startup & VC-Backed Jobs

Engineering Lab Infrastructure Lead

Graphcore - Bristol, England, United Kingdom - In-office - posted 2026-05-08

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Graphcore is building the future of AI compute as part of the SoftBank Group. The company develops the complete AI compute stack—from silicon and software to datacenter-scale infrastructure—and is expanding teams globally to solve critical challenges in artificial intelligence. You will lead the Engineering Lab Support function, ensuring labs, infrastructure, and services remain reliable, scalable, and effective for engineers developing next-generation technologies. This is initially a highly hands-on role where you'll work directly with engineering teams supporting lab environments, managing infrastructure, automating workflows, and solving complex technical problems. As the function grows, you'll build and lead a small team while defining processes, standards, and service models. Key responsibilities include: - Leading development of engineering lab infrastructure and support services - Acting as technical escalation point for complex infrastructure and lab issues - Managing Linux-based servers and engineering environments - Supporting hardware bring-up, validation, and testing activities - Designing and improving operational processes, tooling, and automation - Maintaining and improving network, storage, and compute infrastructure - Managing infrastructure through configuration management and Infrastructure-as-Code - Building strong relationships with engineering teams - Developing knowledge bases, documentation, and operational standards - Recruiting, mentoring, and leading lab infrastructure engineers as the function scales Essential qualifications include strong Linux systems administration experience (Debian/Red Hat), experience supporting engineering/research/HPC/data-center environments, solid networking knowledge (routing, VLANs, VPNs), experience managing physical infrastructure (servers, BMCs, firmware), configuration management tools (Ansible, Puppet), authentication systems (LDAP, RADIUS, Active Directory), strong troubleshooting skills, and excellent communication. Desirable skills include container technologies (Docker, Kubernetes), monitoring platforms (Prometheus, Grafana), Python scripting, CI/CD tooling (GitLab, GitHub Actions), hardware development/silicon validation experience, performance analysis, and web infrastructure technologies.

Similar roles