SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Engineer, Datacenter Server Lifecycle

Anthropic - San Francisco, CA, United States - Hybrid - posted 2026-03-11

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 320,000 - 405,000 / annual

Anthropic is investing $50 billion in American computing infrastructure, including custom-built datacenters optimized for frontier AI model development and serving. This Staff Engineer role sits at the heart of that effort, owning the end-to-end operational journey of every machine in the datacenter fleet—from initial provisioning and deployment through steady-state operation, maintenance, repair, and decommissioning. This is greenfield work where you will help define the services, tooling, standards, and processes that govern how Anthropic operates critical hardware at scale. You'll define AI-native workflows for datacenter operations and drive innovations in efficiency, performance, and reliability. A core distinguishing aspect is the deep intersection with security. The machines handle sensitive workloads including frontier model training and serving millions of Claude users. You will ensure every machine in the fleet is trusted, attested, and operating with a verified chain of integrity from hardware upward. You'll partner closely with the Infrastructure Security team to define and enforce trusted compute standards across the entire lifecycle, from secure provisioning through end-of-life handling. Key responsibilities include: building automation to support datacenter fleets at scale; defining and owning the end-to-end system lifecycle strategy with automation and operational procedures for common events (hardware failures, firmware upgrades, fleet rotations); partnering with Infrastructure Security on trusted compute standards; collaborating with the Networking team on end-to-end connectivity; and building tooling to track machine health, configuration, and operational status across the full fleet. Required qualifications: hands-on experience with server hardware (rack deployment, cabling, troubleshooting, failure modes at scale); end-to-end understanding of hardware lifecycle management (asset tracking, provisioning, maintenance, decommissioning); proficiency in at least one programming language (Python, Rust, Go, Java); working knowledge of modern cloud infrastructure (Kubernetes, AWS/Azure/GCP); ability to communicate and build consensus across stakeholders; comfort with ambiguity and complex cross-functional problems; willingness to travel occasionally to North American datacenter sites. Preferred: 8+ years in datacenter infrastructure management; hands-on experience with GPU/AI accelerator hardware (NVIDIA A100/H100, Google TPUs, AWS Trainium); familiarity with LinuxBoot and NixOS; experience building datacenter automation or fleet management platforms; experience deploying OS distributions across large server fleets; background in capacity planning and hardware refresh strategy at hyperscalers; experience with trusted compute and hardware security (secure boot, TPM, attestation, firmware verification).

Similar roles