SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Snowflake is seeking a Senior Software Engineer to join the Capacity Engineering team, which is responsible for provisioning and optimizing cloud resources across AWS, Azure, and GCP. This role focuses on building a centralized, self-serve internal capacity platform that models demand, forecasts requirements, and delivers optimal CPU and GPU capacity while maximizing fleet utilization and driving hardware cost-efficiency.
You will design and build the capacity platform that unifies CPU and GPU allocation, procurement, reservation lifecycle, and utilization across all three major cloud providers. A key responsibility is owning the canonical capacity data layer—ingesting and reconciling demand forecasts, provider supply signals, commitments, and fleet utilization into a single, trustworthy model consumed company-wide.
The role involves serving as a liaison to Cloud Service Providers, managing and integrating vendor relationships into the capacity planning and procurement workflow. You will build planning and allocation systems that translate demand into hardware requirements (shape, quantity, region, timing) and surface supply risk early with real-time visibility into fleet and reservation health.
You'll drive efficiency by instrumenting utilization across CPU and GPU accelerator workloads, establishing price/performance baselines, and building tooling that recovers stranded capacity and right-sizes commitments. The platform must integrate hardware evolution, evaluating new CPU and GPU generations and their price/performance while building flexibility for backup and cross-family fallbacks.
Partnering with core services, warehouse, AI/ML, and finance teams, you'll forecast and procure capacity ahead of launches, support AI/ML workloads reliably, and translate insights into procurement and allocation decisions. You'll also ensure high availability, reliability, and performance of capacity systems through on-call rotations and incident management.
Required qualifications include 7+ years of industry experience designing, building, and supporting large-scale production systems; hands-on experience with cloud providers on compute cluster and cloud services provisioning (CPU and/or GPU fleets); experience with capacity planning, procurement, resource management, or efficiency work on large private or public cloud systems; deep system and architectural analysis experience; proficiency in Go, Python, or Java; excellent problem-solving and troubleshooting skills; strong communication and collaboration abilities; and a BS/MS in Computer Science, Engineering, or related field.