SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: GBP 325,000 - 485,000 / annual
Anthropic is seeking a Senior Staff Software Engineer to lead the Node Infrastructure team, which owns the full lifecycle of accelerator capacity across Anthropic's AI compute fleet. This role is central to Anthropic's mission of building reliable, interpretable, and steerable AI systems.
You will own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair across GPUs, TPUs, and Trainium accelerators. You'll drive cross-team initiatives to build and scale AI clusters across multiple cloud providers and custom datacenters, designing systems that automatically detect, isolate, and remediate unhealthy hardware to maximize fleet uptime and minimize stranded capacity.
Key responsibilities include defining infrastructure architecture for the hardest problems, working closely with cloud providers and internal research/inference/product teams to shape long-term compute strategy, establishing operational excellence practices (incident response, postmortem culture, on-call), and mentoring engineers through technical leadership.
You should have deep expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, AWS/GCP/Azure), strong proficiency in systems languages (Rust, Go, Python) and Terraform, and hands-on experience with ML accelerators. A track record of leading complex, multi-quarter technical initiatives spanning multiple teams is essential.
Preferred qualifications include 12+ years of software engineering experience with time as a technical lead, experience managing large-scale compute infrastructure (10K+ nodes), depth in Kubernetes internals or cluster orchestration systems, low-level systems experience (kernel, virtualization, device drivers), familiarity with high-performance networking (EFA, RDMA, InfiniBand), and contributions to relevant open-source projects.