SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Perplexity AI - San Francisco, CA, United States - In-office - posted 2026-08-05

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Perplexity is hiring a Member of Technical Staff to own the design and operation of its GPU cluster infrastructure platform. The role sits at the intersection of infrastructure engineering and ML systems, responsible for building a self-serve compute platform that abstracts away the complexity of managing GPUs across multiple cloud providers (CoreWeave, AWS, GCP, etc.). Key responsibilities include: - Design and build a self-serve platform enabling inference engineers and researchers to launch training jobs and operate inference services without managing GPU provisioning or provider-specific details. - Own the full lifecycle of the GPU fleet: provisioning, reliability, capacity management, and integration across multiple cloud providers. - Develop scheduling and placement logic to efficiently allocate scarce GPU capacity across providers and workloads. - Support dual workload patterns: long-running distributed training jobs and high-availability production inference services on the same infrastructure. - Own Kubernetes orchestration for GPU workloads, including custom operators, CRDs, and multi-cluster federation across providers. - Build fault tolerance, autoscaling, and observability systems that keep the fleet utilized and resilient to node failures and provider disruptions. - Set technical direction across infrastructure and inference teams, translating operational constraints into coherent platform architecture. Required depth includes: advanced Kubernetes (custom operators, CRDs, multi-cluster federation), GPU cluster management at scale (NVIDIA hardware, CUDA, high-speed interconnects like InfiniBand/RoCE), multi-cloud orchestration experience, distributed systems fundamentals (scheduling, resource allocation, fault tolerance), and systems-level programming in Go, Rust, or C++. You should have hands-on experience supporting both training and inference workloads and understand their opposing infrastructure demands. Valued additional experience: inference serving stacks (vLLM, SGLang, TensorRT-LLM), HPC schedulers (Slurm), GPU kernel work (CUDA/Triton), production high-speed interconnects, and ML observability tools (Prometheus, Grafana, Weights & Biases).

About Perplexity AI

AI / Data / Infrastructure — AI answer engine and search product for consumers and enterprises.

Similar roles