SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Anthropic is seeking a Data Center Operations Lead to manage partner-operated data center sites. This role bridges Anthropic's engineering needs with third-party site operators, ensuring compute fleet availability, deployment velocity, and incident response excellence.
You will own operational outcomes for assigned sites, including availability metrics, deployment milestones, and repair turnaround times. Rather than directly managing operations staff, you provide tactical direction to vendor teams, set priorities, define operational standards, and oversee performance against SLAs. You'll build and refine operational playbooks—both at individual sites and fleet-wide—establishing processes for deployment, break-fix, change management, security, and EHS compliance.
Key responsibilities include: leading weekly operations reviews and scorecards with vendor leads; directing deployment surges to meet compute milestones; analyzing failure patterns and driving root-cause fixes; creating break-fix ownership matrices and training vendor teams; serving as Incident Commander for facility events and producing post-mortems; establishing operational readiness for new data halls; and identifying process gaps to codify improvements as program standards.
You'll participate in an on-call incident escalation rotation and support non-standard hours during deployment surges and maintenance windows. You translate engineering requirements into vendor direction and communicate site constraints and risks to leadership.
Required: 8+ years in data center operations (hardware, IT infrastructure, or critical facilities) with accountability for production availability; vendor/MSP management experience with measurable outcomes (SOWs, SLAs, reviews, corrective action); hands-on technical depth in server, network, and rack-level infrastructure to independently verify vendor claims; experience building or substantially improving operational processes; incident command or lead-responder experience; and a Bachelor's degree in a relevant domain or equivalent practical experience.
Desirable: experience with third-party colocation or partner-operated sites; standing up new data halls from commissioning through first deployment; GPU/accelerator or high-density liquid-cooled infrastructure; and multi-vendor site operations.