SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff — Training

RadixArk - Palo Alto, CA, United States - In-office - posted 2026-02-17

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will work on large-scale distributed training infrastructure for LLMs and generative models, pushing the limits of scale, efficiency, accuracy, and reliability across 10,000 or 100,000+ GPUs. This role sits at the intersection of ML, systems, and performance engineering. Key responsibilities include contributing to open-source large-scale post-training infrastructure (Miles) and inference systems (SGLang), optimizing throughput, scalability, and hardware efficiency, and improving reliability and fault tolerance for long-running training jobs. You will develop training frameworks and infrastructure tooling, collaborate with model researchers to support frontier experiments, debug and resolve cross-layer performance bottlenecks, and build observability systems for training performance and reliability. Required qualifications include 3+ years of experience in ML systems or large-scale training infrastructure, experience building or operating large-scale agentic post-training systems, experience working on training/inference correctness or precision-related problems, experience debugging performance and stability issues in large post-training jobs, and experience improving training or inference efficiency. Strong plus qualifications include experience training 100+ billion-parameter models, experience with train/inference optimization for large-scale RL or other production workloads, familiarity with training stacks (Megatron-LM, FSDP, torchtitan, etc.) and inference stacks (SGLang, vLLM, etc.), familiarity with post-training frameworks (Miles, Slime, veRL, Prime-RL, AReaL, etc.), experience with RDMA, InfiniBand, NVLink, NCCL/RCCL, or high-speed GPU interconnects, contributions to ML systems open-source projects, and experience with checkpointing, fault recovery, and elastic training.

Similar roles