SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer, Compute Foundations

OpenAI - San Francisco, CA, United States - In-office - posted 2026-09-16

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

OpenAI's Compute Foundations team builds the software infrastructure that manages GPU compute across multiple data centers and infrastructure providers, supporting both model training and inference at scale. The team develops Kubernetes-based control planes, controllers, services, and APIs that coordinate machine and cluster lifecycles, bridging global infrastructure management with bare-metal systems realities. In this role, you will design and build distributed systems that provision, configure, and manage compute infrastructure throughout its lifecycle. You'll work on Kubernetes controllers and services that coordinate infrastructure across sites, define APIs and resource models for lifecycle operations, and build provisioning and configuration services that integrate network boot, hardware management, and firmware deployment. Your responsibilities include developing lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning; designing reliable reconciliation and recovery mechanisms that handle concurrent changes and partial failures; and improving control-plane throughput and API latency while respecting provider rate limits. You will partner with hardware, networking, and data-center teams to integrate new sites and GPU hardware generations into the platform. The role combines software architecture with practical understanding of how machines and data centers operate. You'll diagnose reliability and performance problems across service, OS, and machine boundaries, turning production evidence into lasting improvements. QUALIFICATIONS: - Strong software engineering fundamentals with experience designing, implementing, and owning production distributed systems or infrastructure services - Experience developing infrastructure systems using Kubernetes APIs and reconciliation patterns - Understanding of bare-metal node provisioning lifecycle, with depth in areas such as PXE, DHCP/DNS, baseboard management controllers (BMCs), firmware, Linux, drivers, images, or configuration management - Ability to design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures across service and provider boundaries - Capability to diagnose reliability and performance problems across service, OS, and machine boundaries - Effective cross-functional communication and ability to work across engineering specialties BONUS: - Experience building infrastructure control planes coordinating operations across multiple sites or regions - Background with GPU or HPC infrastructure, including topology and shared dependencies - Experience integrating multiple hardware platforms or providers into a unified service model

Similar roles