SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer, Compute Foundations Systems

OpenAI - San Francisco, CA, United States - In-office - posted 2026-07-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

OpenAI's Frontier Systems Foundations team is seeking a systems software engineer to build and maintain the Linux operating-system foundation for the company's frontier GPU compute fleet. This role sits at the intersection of hardware, infrastructure, and AI systems—you'll own the systems software closest to the machine: OS images, kernels, drivers, firmware integration, provisioning, and fleet-wide validation. You will work directly with hardware engineers, vendors, and infrastructure teams to bring new compute platforms online, integrate system components, and debug failures across firmware, boot, operating systems, kernels, drivers, and workload interactions. Your work directly influences how quickly new capacity becomes usable and how reliably large-scale GPU clusters operate at the frontier of AI model training. Key responsibilities include building and maintaining the Linux host software stack for large GPU clusters (Ubuntu/Debian images, kernel configuration, drivers, packages, disks, machine configuration); designing reproducible OS-image builds and package workflows for heterogeneous bare-metal and cloud fleets; integrating and qualifying kernels, modules, drivers, and firmware across new hardware platforms with safe canary and rollback paths; bringing up new hardware platforms and compute SKUs in collaboration with hardware teams; debugging complex system failures that cross multiple layers (boot, provisioning, firmware, disks, kernels, drivers); improving fleet-scale provisioning and maintenance by eliminating manual intervention; and building systems tooling and diagnostics for validation and troubleshooting. Ideal candidates have significant production Linux experience (particularly Ubuntu/Debian), deep expertise in one or more of: Linux kernels/modules/drivers, distribution/package/OS-image engineering, boot/provisioning/disks, or firmware/driver integration. You should be comfortable writing production-quality systems software and automation, understand how to safely deliver system-level changes with compatibility testing and staged rollout, and can systematically debug failures crossing hardware/firmware/OS boundaries. Bonus experience includes bare-metal provisioning tooling, GPU/HPC/accelerator systems, bringing up new server platforms, open-source systems software contributions, or supporting system changes across large production fleets.

Similar roles