SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI is seeking a Software Engineer to build and optimize the inference stack for AWS Trainium accelerators. This is a deeply technical, cross-stack role spanning kernels, compilers, and model execution systems.
You will develop high-performance kernels for critical model operations, extend compiler support to efficiently target Trainium hardware, and build systems to execute and optimize model forward passes. The role involves profiling workloads to identify bottlenecks across kernels, compiler-generated code, runtime, and model execution layers. You'll partner closely with inference and ML systems teams to bring new models and architectures onto Trainium, working across the hardware/software boundary to unlock performance from specialized AI accelerators.
Key responsibilities include owning complex performance and systems problems end-to-end from investigation through production deployment, reasoning about performance across multiple abstraction layers, and learning new hardware and software domains as needed.
Required qualifications: 3+ years of relevant engineering experience in ML systems, compilers, kernels, runtimes, or performance engineering. Strong systems programming fundamentals and experience writing performance-critical software. Experience working with GPU, TPU, Trainium, or other specialized accelerator architectures. Ability to reason about performance across multiple stack layers from hardware through ML frameworks.
Bonus qualifications include AWS Trainium or AWS Neuron SDK experience, contributions to ML frameworks like PyTorch or JAX, compiler infrastructure work (LLVM, MLIR, XLA, Triton), or kernel development for specialized accelerators.
OpenAI is an AI research and deployment company dedicated to ensuring general-purpose AI benefits all of humanity, with a focus on safety and responsible deployment.