SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's Industrial Compute organization is seeking a Technical Program Manager to lead capacity planning across large-scale AI infrastructure. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations.
You will own end-to-end capacity planning processes spanning near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Key responsibilities include translating research, training, inference, and product demand into concrete compute, accelerator, cluster, networking, storage, rack, power, and site requirements. You will develop scenarios that make assumptions, confidence levels, constraints, and decision points explicit, then reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity.
This is not a finance-only forecasting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change rapidly. You will partner with research and engineering teams to understand workload priorities, technical dependencies, and utilization patterns, while coordinating with sourcing, finance, hardware, deployment, and operations teams on supply commitments, activation schedules, costs, and delivery risks.
You will build dashboards, analytical tools, and executive updates that create a credible source of truth for infrastructure capacity. You will track infrastructure lead times, critical dependencies, utilization, headroom, forecast accuracy, activation readiness, and capacity risk. When infrastructure is constrained or delivery plans change, you will support allocation and prioritization decisions and drive mitigation plans for site delays, hardware shortages, network or storage constraints, and other capacity risks. As OpenAI's infrastructure scales, you will continuously improve planning models, governance, data quality, and accountability.
The role requires regular in-person collaboration with infrastructure, research, and operational partners in San Francisco.
QUALIFICATIONS:
- 8+ years of experience in technical program management, infrastructure capacity planning, cloud infrastructure, supply planning, or closely related field
- Demonstrated ownership of capacity planning for large-scale distributed systems, cloud platforms, AI/ML workloads, or hyperscale infrastructure
- Ability to translate ambiguous demand into structured assumptions, scenarios, technical resource requirements, and executable plans
- Technical fluency across compute, networking, storage, data center infrastructure, utilization, reliability, and deployment dependencies
- Strong analytical skills and experience developing planning models, operational metrics, dashboards, or data-driven decision systems
- Experience leading decisions across engineering, finance, sourcing, deployment, operations, and executive stakeholders
- Excellent written and verbal communication, including ability to explain uncertainty, tradeoffs, and recommendations clearly
PREFERRED:
- Experience planning GPU, accelerator, or AI infrastructure capacity for training or inference workloads
- Experience with cluster allocation, cloud capacity, hardware supply, infrastructure procurement, site readiness, or production activation
- Familiarity with forecasting methods, scenario planning, optimization, cost attribution, capacity economics, and utilization management
- Experience building planning tools or source-of-truth systems using SQL, Python, spreadsheets, BI platforms, or similar technologies
- Track record of improving utilization, reducing infrastructure cost or risk, and enabling critical workloads during periods of constrained supply