SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer - Data Aquisition (systems)

OpenAI - San Francisco, CA, United States - In-office - posted 2026-07-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

OpenAI is seeking a Software Engineer to join the infrastructure team that builds and operates systems enabling researchers to run reliable, scalable, and efficient research workflows. This role sits at the intersection of software engineering, infrastructure, systems administration, and reliability engineering. You will build and operate infrastructure supporting frontier research and critical research-facing systems, working close to the metal while maintaining a software-engineering mindset. The work spans networking, Kubernetes, cluster operations, bootstrapping, provisioning, automation, and reliability. You'll support systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and cluster-wide services. Key responsibilities include: building and operating reliable infrastructure for research workloads; improving cluster bootstrapping, provisioning, and deployment workflows; debugging issues across networking, compute, storage, orchestration, and service reliability; writing software and automation to reduce manual operational work; partnering with researchers and infrastructure engineers to translate system needs into durable solutions; and driving critical systems independently from problem definition through execution. Ideal candidates have strong systems fundamentals, understand how infrastructure scales in practice, are comfortable with Linux, networking, Kubernetes, and distributed systems operations, and can write software to automate and improve infrastructure. You should have a strong execution mindset, enjoy supporting a wide surface area of systems, and be pragmatic about building custom systems versus using existing tools. Nice-to-have experience includes PXE boot, bare-metal infrastructure, large-scale fleet management, operating Kubernetes at scale, infrastructure-as-code, CI/CD, observability, deployment automation, search infrastructure, data platforms, and environments where reliability, scale, and speed all matter.

Similar roles