SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff — Cluster Infrastructure

RadixArk - Palo Alto, CA, United States - In-office - posted 2026-02-17

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

RadixArk is seeking a Member of Technical Staff for Cluster Infrastructure to architect and scale the core compute platform powering frontier-level AI training and inference. This is a deep systems engineering role focused on designing and operating highly reliable, high-performance GPU/TPU clusters, building next-generation scheduling and resource management systems, and optimizing large-scale distributed infrastructure for AI workloads. Key responsibilities include architecting and scaling large AI compute clusters for training and inference, designing cluster management and scheduling systems, optimizing performance and utilization of GPU/TPU infrastructure, improving fault tolerance and system resilience at scale, driving observability and monitoring for cluster infrastructure, and collaborating with ML and systems engineers to support production AI workloads. Required qualifications: 5+ years in distributed systems, infrastructure, or large-scale compute platforms; strong background in distributed systems design and architecture; deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers); hands-on production experience with GPU/TPU infrastructure; strong Linux systems and networking fundamentals; proficiency in Go, Rust, C++, or Python for production systems; experience debugging complex multi-layer issues across hardware, OS, networking, and distributed services; proven ability to design reliable, scalable systems in production. Strong plus factors include experience with large-scale ML/AI workloads, familiarity with RDMA/InfiniBand or high-performance networking, experience operating clusters at 1000+ GPU scale, background in HPC or performance-critical systems, and open-source contributions in systems or infrastructure.

Similar roles