SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 320,000 - 405,000 / annual
Anthropic is seeking a Repairs Program Lead to define and manage the end-to-end hardware repair program across its growing fleet of data centers. In this role, you will be accountable for repair turnaround time and compute returned to service across every site, covering server, GPU/accelerator, network, and optics break-fix, RMA and reverse logistics with OEMs and ODMs, and spares/repair inventory management.
Key responsibilities include: defining global repair strategy with SLAs, prioritization rules, and escalation paths; owning repair turnaround time and backlog across the fleet via dashboards; authoring and improving procedures for triage, break-fix, and return-to-service validation; managing RMA and reverse logistics programs with vendors; setting spares pool sizing and stocking levels by site and part; analyzing failure patterns to identify root causes and drive corrective actions; leading operating cadence with vendor and site leads including weekly reviews and scorecards; and communicating repair constraints and fleet availability impact to engineering and leadership.
You should have 8+ years of data center operations experience as a manager or technical lead with accountability for production availability. Demonstrated track record running break-fix programs at large scale across multiple sites is essential. You must have managed vendors, OEMs, or contract workforces to measurable outcomes (SLAs, operational reviews, corrective action). Hands-on technical depth in server, network, and rack-level hardware is required to independently verify repair quality. Experience building or substantially improving operational processes, working with ticket/telemetry/inventory data, and a bachelor's degree in relevant domain (or equivalent) are expected.
Bonus qualifications include experience with GPU/accelerator or high-density liquid-cooled infrastructure; managing RMA and warranty programs with hyperscale OEMs/ODMs; spares planning, reverse logistics, or depot repair at data center scale; delivering repair outcomes in partner-operated or colocation sites with third-party staffing; and familiarity with optics and high-speed interconnect failures.