SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's Hardware organization is building AI-native silicon and system-level solutions for advanced AI workloads. This role focuses on developing the model runtime within the inference engine that executes frontier models at scale on OpenAI's custom silicon.
You will design and implement the LLM inference runtime, building scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Key responsibilities include developing distributed execution strategies across chips, hosts, and racks; optimizing end-to-end latency, throughput, memory efficiency, and hardware utilization; and partnering with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks.
You will create profiling, observability, benchmarking, and performance-modeling tools to make runtime behavior measurable and actionable. This includes debugging complex correctness, performance, and reliability issues spanning model code, runtime software, communication layers, and hardware, then translating workload insights into requirements for future silicon generations.
Ideal candidates have strong systems programming experience in C++, Rust, Python, or comparable performance-oriented languages. You should have built or optimized runtimes, distributed systems, compilers, kernels, or model-serving infrastructure. Deep understanding of modern LLM inference (prefill/decode behavior, batching, KV-cache tradeoffs, model parallelism) is essential. You can reason quantitatively about latency, throughput, compute intensity, memory bandwidth, and utilization. You're comfortable profiling and debugging across multiple hardware-software stack layers, designing clean abstractions while retaining low-level control for specialized hardware performance extraction. You work effectively across cross-functional teams and prioritize production quality including correctness, observability, reliability, and maintainability at scale.