SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
RadixArk, an infrastructure-first AI company founded by veterans from xAI and NVIDIA, is seeking a Member of Technical Staff to build high-performance inference and training systems on Google TPU hardware. You will work on critical infrastructure including SGLang-JAX, optimizing large-model workloads to maximize efficiency on the latest tensor processing units.
In this role, you will design and implement high-performance systems using JAX, XLA, and Pallas, pushing model workloads to their limits on TPU hardware. Key responsibilities include optimizing end-to-end latency and throughput for LLM serving, designing SPMD strategies for distributed inference and training, implementing custom Pallas kernels for performance-critical operations, and profiling XLA compilation pipelines and HLO graph transformations. You will collaborate with kernel engineers and compiler teams to achieve performance wins across the full stack and contribute to open-source projects with optimization guides and benchmarks.
Required qualifications include 3+ years building production ML systems with JAX/Torch, XLA, or TPU-focused frameworks, and a Bachelor's or Master's degree in Computer Science, Electrical Engineering, or equivalent industry experience. You should have deep understanding of XLA internals (HLO, MLIR, operator fusion, SPMD partitioning), strong performance tuning instincts across compiler and runtime layers, and experience with distributed inference systems like SGLang or vLLM. Proficiency in Python with demonstrated ability to write high-performance production code is essential. Experience writing custom GPU/TPU/AI accelerator kernels and familiarity with Pallas for kernel development is strongly preferred.