SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff — Inference

RadixArk - Palo Alto, CA, United States - In-office - posted 2026-02-17

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization. Your work will directly shape how state-of-the-art models are deployed and experienced by users worldwide. This is a deeply technical, high-impact role for engineers who enjoy working close to the hardware–software boundary and solving performance-critical problems at scale. Key responsibilities include designing and building large-scale inference systems for frontier AI models, optimizing latency, throughput, and GPU utilization in production inference, developing and improving model serving architectures and runtimes, and working on batching, scheduling, and memory management strategies. Required qualifications: 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems; strong expertise in large-scale inference systems for LLMs or generative models; deep understanding of GPU architecture and performance characteristics; experience optimizing latency- and throughput-critical production systems; strong knowledge of distributed systems and networking fundamentals; proficiency in Python, Rust, C++, or Go for production systems; experience profiling and optimizing compute-intensive workloads; and strong debugging skills across system layers (model, runtime, kernel, network). Strong plus qualifications include experience with LLM serving stacks (SGLang, vLLM, TensorRT-LLM), familiarity with CUDA, Triton, or custom kernel optimization, experience with batching, KV-cache management, and scheduling strategies, experience running inference at scale (1000+ GPUs), background in HPC or high-performance systems, and open-source contributions in ML or systems infrastructure.

Similar roles