SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Staff+ Software Engineer, Node Infra

Anthropic - San Francisco, CA, United States - Hybrid - posted 2026-04-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 405,000 - 485,000 / annual

Anthropic is seeking a Senior Staff Software Engineer to lead the Node Infrastructure team, which owns the full lifecycle of accelerator capacity across Anthropic's AI compute fleet. This role is foundational to Anthropic's mission of developing reliable, interpretable, and steerable AI systems. You will own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair across GPUs, TPUs, and Trainium accelerators. Key responsibilities include driving cross-team initiatives to build and scale AI clusters across multiple cloud providers, designing systems that automatically detect and remediate unhealthy hardware to maximize fleet uptime, and defining infrastructure architecture for some of the industry's hardest problems. You'll work closely with cloud providers and internal research, inference, and product teams to shape long-term compute and infrastructure strategy. You'll establish operational excellence practices including incident response and on-call culture, and mentor engineers on your team through technical coaching. Required qualifications include deep expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, AWS/GCP/Azure); strong proficiency in systems languages like Rust, Go, or Python with Terraform; hands-on experience with ML accelerators; and a track record of leading complex, multi-quarter technical initiatives across teams. You should be able to build alignment with senior stakeholders and communicate effectively at all levels. Preferred experience includes 12+ years of software engineering with time as a technical lead, managing large-scale compute infrastructure (10K+ nodes), depth in Kubernetes internals or cluster orchestration systems, low-level systems experience (kernel, virtualization, device drivers), familiarity with high-performance networking for distributed ML, and contributions to relevant open-source projects.

Similar roles