SlipstreamJobsFresh Startup & VC-Backed Jobs

Simulation Infrastructure Engineer

OpenAI - San Francisco, CA, United States - Hybrid - posted 2026-07-22

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

OpenAI's Robotics team is seeking a Simulation Infrastructure Engineer to build production-quality simulation pipelines that power model training, evaluation, and hardware-in-the-loop validation. This role owns the automation, orchestration, and tool integration that apply simulation to concrete robotics tasks. Key responsibilities include: - Build and maintain CI/CD pipelines for simulation code, environments, and tasks, ensuring simulation artifacts are testable, versioned, and reproducible. - Implement end-to-end automation to run model evaluation in simulation (SIL) and orchestrate hardware-in-the-loop (HIL) runs; compute realism and task metrics, generate dashboards and alerts. - Create robust APIs and connectors so research, training, and data-collection systems can schedule, seed, and evaluate batches of simulations; support RL rollouts and imitation-data collection. - Build scheduling, batching, and orchestration for running tens of thousands of concurrent rollouts and large RL workloads; solve engine-level scaling and optimize cloud/GPU runtime reliability. - Produce metrics and tooling for measuring simulation health, throughput, fidelity regressions, and cost; create presubmit and canary tests that catch regressions early. - Implement artifact versioning, environment immutability, experiment provenance, and policies for resource quotas and cost control. - Work closely with Sim Environments, Sim Realism, research, and ops teams to ensure simulation improvements translate into better model evaluation and training results. Ideal candidates have deep software engineering and infrastructure experience, including CI/CD at scale, reliable pipeline authoring, and production services that coordinate many moving parts. Comfort with distributed compute, cloud GPU workloads, and concurrent simulation scheduling is essential. Experience with HIL/SIL workflows, hardware-software integration, automation, observability, and metrics is highly valued. Strong proficiency in Python, C++, or Rust; container orchestration (Kubernetes); distributed task queues; and CI systems is required. Experience with RL tooling, task generators, or large-scale data pipelines is a bonus. The role requires 4 days per week in-person presence in San Francisco.

Similar roles