SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
RadixArk is seeking a Member of Technical Staff focused on Accelerator Systems to optimize performance for frontier AI infrastructure. In this role, you will bring up, optimize, and maintain SGLang, Miles, and RadixArk's infrastructure stack across NVIDIA and AMD GPUs, Google TPUs, modern server CPUs, and emerging AI accelerators. You'll port kernels and runtimes onto unfamiliar hardware, design abstractions that keep a single codebase fast across all platforms, and work directly with silicon partners on pre-release hardware with immature tooling.
This is one of the broadest technical roles at RadixArk. You'll tackle evolving performance challenges: memory hierarchies that break previous assumptions, compilers that fuse differently, and collective libraries that don't yet exist. The ideal candidate finds this variety appealing and can go deep on new architectures quickly without losing performance instincts built on prior platforms.
Key responsibilities include: bringing up inference and training systems on new accelerator platforms and driving them to competitive performance; designing hardware abstractions that enable a single codebase to stay fast across vendors without forking; porting and optimizing kernels across programming models and memory architectures; building cross-platform benchmarking, profiling, and regression detection; debugging numerical divergence and correctness gaps between platforms; collaborating with vendor engineering teams on pre-release hardware and compiler issues; partnering with kernel, runtime, distributed systems, and product engineers; serving as the internal source of truth on platform capabilities; and contributing hardware-specific optimizations back to open-source SGLang and Miles.
Required: 4+ years in systems, performance, or ML infrastructure engineering; deep expertise in at least one accelerator programming model (CUDA, ROCm/HIP, Pallas/XLA, Triton, or vendor SDK) with ability to pick up new ones quickly; strong understanding of accelerator architecture (memory hierarchy, bandwidth limits, occupancy, tradeoffs); experience writing or optimizing high-performance kernels for ML workloads; experience with distributed execution and communication libraries (NCCL, RCCL, MPI); proficiency in C++ and Python; strong debugging and profiling skills at system level; track record of performance work shipped to production.
Strong pluses include: experience bringing up ML workloads on new silicon; hands-on depth in multiple vendor ecosystems; compiler stack experience (XLA, MLIR, TVM, Triton); hardware abstraction layer or portable kernel interface design; quantization and mixed-precision work; distributed inference systems (SGLang, vLLM) or training frameworks (Miles, Megatron, veRL, TorchTitan); CPU inference optimization; collective communication optimization at scale; open-source kernel/compiler/ML systems contributions; direct vendor or cloud partner collaboration; HPC or performance-critical systems background.