SlipstreamJobsFresh Startup & VC-Backed Jobs

Performance Engineer, Inference Engine

Anthropic - San Francisco, CA, United States - Hybrid - posted 2026-09-09

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 350,000 - 850,000 / annual

Anthropic is seeking a Performance Engineer to optimize and build the inference engine—the core software layer that manages token flow between accelerator kernels and routing layers. This system handles request batching, model distribution across chips, memory management for weights and activations, forward pass coordination, and model state across requests. It runs on all of Anthropic's accelerator platforms, serving Claude to millions of users and powering research workloads. You will work at Anthropic scale to improve throughput, cost, reliability, and latency across all accelerator and cloud platforms. The role requires deep familiarity with hardware performance characteristics (FLOPs, HBM, PCIe, RDMA, network links) and the ability to quickly model where time and bytes are spent and what constrains performance. This is a deeply technical, high-impact position suited for engineers comfortable working across accelerator programming, high-performance host-device coordination, and large-scale distributed systems. Key responsibilities include keeping device utilization high (accelerators should never idle due to overhead), reusing model state instead of recomputing, measuring and modeling performance gaps before implementing changes, ensuring model quality across platforms, and supporting production safety systems with efficient inference infrastructure. Minimum qualifications include a working mental model of LLM inference (prefill/decode on accelerators, memory, interconnect, and host coordination), proven ability to ramp quickly on unfamiliar systems and ship consequential changes, strong systems programming in Rust/C++, analytical performance debugging (profile, hypothesize, test, measure), low ego and collaborative mindset, and enjoyment of pair programming. Preferred qualifications include experience inside LLM serving engines, GPU/accelerator programming, OS internals, language modeling with transformers, experience building allocators/caches/schedulers/high-bandwidth transports, Rust fluency, and experience with reproducible systems (determinism, replay, property-based testing).

Similar roles