SlipstreamJobsFresh Startup & VC-Backed Jobs

Director of Data Center Facilities - AI Infrastructure

TensorWave - Remote - Remote

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

TensorWave is seeking a Director of Data Center Facilities - AI Infrastructure to lead the physical infrastructure strategy, operations, reliability, and expansion of high-density AI data centers built for GPU computing at scale. This is a highly visible leadership role overseeing mission-critical facilities that support breakthrough AI compute platforms. Key Responsibilities: AI Data Center Operations: Provide overall leadership for facilities operations supporting high-density AI/GPU compute environments. Ensure consistent delivery of power, cooling, environmental conditions, and infrastructure availability. Establish operational readiness standards for new AI compute deployments and lead facility response to electrical, mechanical, thermal, and cooling incidents. High-Density Power Infrastructure: Oversee utility service, substations, medium-voltage distribution, transformers, switchgear, UPS systems, generators, busways, PDUs, and rack-level power distribution. Partner with utilities and engineering teams on capacity, grid interconnection, and resiliency. Develop long-range power capacity plans aligned with GPU deployment schedules. Advanced Cooling & Thermal Management: Lead design, operation, and optimization of cooling systems for high-density AI compute. Establish standards for liquid-cooling reliability, leak detection, water quality, and maintenance. Partner with hardware teams to understand evolving thermal requirements and develop strategies for increasing rack density without compromising reliability. AI Capacity Planning & Infrastructure Strategy: Develop multi-year facilities capacity plans based on AI compute roadmaps. Translate compute requirements into MW, cooling tonnage, rack density, and space requirements. Participate in site selection, utility strategy, and campus master planning. Evaluate emerging infrastructure technologies for large-scale AI environments. New Construction, Expansion & Commissioning: Lead facilities participation in design, construction, commissioning, and turnover of new AI data center capacity. Establish commissioning requirements and review engineering designs. Develop standardized infrastructure designs replicable across multiple sites. Reliability & Incident Management: Establish world-class reliability programs for AI data center infrastructure. Lead root-cause analysis for critical incidents and implement corrective actions. Develop predictive maintenance strategies and conduct scenario-based exercises for major failures. Automation, Controls & Data: Drive increased use of BMS, EPMS, DCIM, telemetry, and predictive monitoring. Establish real-time visibility into electrical and thermal capacity and identify automation opportunities. Team Leadership: Build and lead high-performing organizations of facilities engineers, managers, technicians, and contractors. Establish structures for 24x7 operations at scale. Develop technical training and succession planning programs. Financial & Vendor Management: Manage strategic relationships with utilities, OEMs, engineering firms, and equipment manufacturers. Establish performance requirements and SLAs. Evaluate lifecycle costs and identify cost-reduction opportunities. Safety, Compliance & Risk: Establish safety-first culture for high-voltage systems and heavy equipment. Develop emergency response and business-continuity programs. Conduct infrastructure risk assessments and ensure audit readiness. Energy & Sustainability: Develop strategies to manage significant energy demands. Optimize PUE, WUE, and carbon impact. Partner on renewable energy and evaluate alternative cooling technologies. Required Qualifications: - Bachelor's degree in Electrical Engineering, Mechanical Engineering, Facilities Engineering, or related technical discipline; equivalent experience considered - 10+ years of experience in mission-critical data centers, high-density computing environments, critical infrastructure, or comparable facilities - 5+ years of experience leading large technical facilities organizations - Demonstrated experience managing large-scale electrical and mechanical infrastructure - Strong understanding of high-density data center power architectures and cooling systems - Experience supporting or operating facilities with high rack power densities and rapidly scaling compute infrastructure - Experience with mission-critical maintenance, reliability engineering, incident management, and operational readiness - Demonstrated experience managing significant operating and capital budgets - Strong understanding of infrastructure capacity planning and facility expansion

Similar roles