SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Fal is a generative media infrastructure platform that powers AI products at scale. The company provides high-performance inference, orchestration, and observability tools for developers and enterprises building AI-native applications.
You will lead all aspects of fal's data center operations, owning the full lifecycle from strategy and planning through design, build-out, deployment, and steady-state management of the GPU-accelerated infrastructure that powers fal's inference services globally.
Key responsibilities include:
- Own the complete data center operations lifecycle across all fal sites, including capacity planning, infrastructure design, facility build-out, and deployment execution.
- Develop and execute infrastructure operations strategy with clear OKRs and KPIs, balancing capacity, cost, reliability, and scale in close coordination with engineering, finance, and leadership.
- Directly manage on-site operations teams including site leads and technicians; own incident management, escalation protocols, and drive uptime and MTTR improvements.
- Partner with engineering, network, and capacity planning teams to execute infrastructure expansion while optimizing power efficiency, cooling systems, rack density, and deployment timelines.
- Build operational processes, documentation standards, and vendor governance frameworks as fal scales its global infrastructure footprint.
You bring 10+ years of data center and infrastructure operations experience, including senior leadership (Director level) managing multi-site or global operations. You have hands-on experience operating large-scale (10MW+) mission-critical data centers and full-lifecycle expertise from planning and build-out through steady-state operations, ideally in high-density GPU or AI inference environments. You are strategic and cross-functional, comfortable partnering with engineering, finance, and business leadership on capacity, cost, and scaling decisions. You have strong commercial and vendor relationship management skills, with experience in colocation and leased-space operations.
Bonus experience includes standing up new data center capacity from the ground up, operating high-density LPU/GPU or liquid-cooled environments, and building and scaling data center operations teams.