SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
San Francisco Compute is building the next generation of GPU cloud infrastructure, enabling customers to lease large-scale supercomputers without catastrophic balance sheet risk. Unlike traditional GPU cloud providers that force long-term contracts, SFC operates a hybrid model combining the technical control of a neocloud with the flexibility of a compute market, allowing customers to sublease capacity.
The company was founded by leaders from Lambda, Crusoe, Digital Ocean, AWS, and Hut8, with a CTO who previously founded and led Voltage Park. The team has deployed 8GW of datacenter capacity and managed hundreds of thousands of GPUs across prior roles.
As an HPC/GPU Cluster Architect, you will be a core member of the infrastructure team responsible for architecting, deploying, and operating GPU clusters globally. You'll participate in on-call rotation, deploy new environments, troubleshoot issues across the full stack, and drive automation to enable deployments at scale. As an early contributor to a small but ambitious team, you'll help shape culture, mentor junior engineers, and work directly with customers to understand their needs.
Key responsibilities include:
- Architecting and deploying new GPU clusters around the world
- Designing and implementing fleet automation (provisioning, monitoring, remediation)
- Debugging performance and reliability issues across hardware, OS, drivers, and networking
- Creating and maintaining operational documentation and runbooks
- Mentoring junior engineers and contributing to team culture
- Participating in on-call rotation and incident response
- Traveling domestically as needed for cluster deployments and operations
Requirements:
- 5+ years of hands-on experience designing, architecting, and scaling at least one HPC or GPU compute cluster in production (ideally >1,000 GPUs, though not required)
- Deep understanding of server hardware fundamentals: GPUs, NICs, PCIe, memory, thermals, and power
- Ability to debug performance and reliability issues across hardware, OS, drivers, and networking layers (full-stack)
- Comfort with infrastructure-as-code and fleet automation
- Strong operational discipline: ability to create and value comprehensive documentation and runbooks
- Willingness and ability to mentor junior engineers
- Openness to domestic travel when required
Nice to haves:
- Data center operations experience (power, cooling, colo/vendor engagements)
- Strong Linux systems administration (kernel drivers, RDMA stack tuning, performance analysis)
- Experience with schedulers and orchestration (Slurm, Kubernetes)
- Virtualization technologies (KVM, QEMU, libvirt)
- Telemetry pipelines for predictive hardware failure detection
- High-speed fabric troubleshooting (InfiniBand, RoCEv2 Ethernet)