SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff+ Software Engineer, Infrastructure (Distributed Systems)

Anthropic - San Francisco, CA, United States - Hybrid - posted 2025-10-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 320,000 - 485,000 / annual

Anthropic is building reliable, interpretable, and steerable AI systems. The Infrastructure organization operates the distributed systems that train, serve, and secure Anthropic's AI models, including data pipelines, some of the world's largest Kubernetes clusters, databases, observability tools, and developer tooling that all teams depend on. As a Staff+ Software Engineer on the Infrastructure team, you will independently scope and lead complex, multi-month infrastructure projects from ambiguous starting points through production deployment. You'll make architectural decisions that shape the foundation other engineers and teams build upon, drive technical alignment across multiple teams, and partner with research and product teams to translate their infrastructure and compute needs into technical designs. Key responsibilities include: independently scoping and leading complex infrastructure projects; making architectural decisions that influence the broader engineering organization; driving alignment on technical direction across teams; partnering with research and product to understand evolving needs; owning reliability, scalability, and security of systems as they scale; setting technical strategy and standards for your team; building and improving operational processes like incident response and on-call rotations; and mentoring other engineers. Minimum qualifications: experience designing, building, and operating large-scale distributed systems in production; track record of independently scoping and delivering complex, multi-month technical projects; experience making architectural decisions others build on; strong software engineering fundamentals and proficiency in Python, Rust, Go, or Java; experience with modern cloud infrastructure including Kubernetes and infrastructure-as-code on AWS/GCP; strong written and verbal communication with cross-team alignment experience. Preferred: 10+ years of software engineering experience; experience with ML infrastructure (GPUs, TPUs, Trainium, NCCL); low-level systems experience (Linux kernel tuning, eBPF); security or privacy engineering background; prior technical leadership or mentoring experience. Team placement occurs after interviews based on interests, experience, and organizational needs. Minimum education: Bachelor's degree or equivalent. Hybrid policy: at least 25% in-office time, though some roles may require more. Visa sponsorship available.

Similar roles