SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, Observability & Profiling

Anthropic - London, United Kingdom - In-office - posted 2026-09-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: GBP 325,000 - 390,000 / annual

Anthropic is seeking a Staff Software Engineer to join the Observability team within Infrastructure. This role owns the monitoring and telemetry infrastructure that powers Anthropic's research and product systems—spanning metrics, logging pipelines, distributed tracing, profiling, error analytics, alerting, and query interfaces. As Anthropic scales across massive GPU, TPU, and Trainium clusters, operational data volume and complexity are growing exponentially. The hardest problems increasingly live below the application layer. You'll design and build next-generation observability systems: high-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing and network visibility, and agentic diagnostic tools. The goal is enabling engineers to detect, diagnose, and resolve issues in minutes rather than hours—even when answers lie in the kernel, network stack, or accelerator. Key responsibilities include: designing scalable telemetry ingest and storage pipelines for metrics, logs, traces, and errors across multi-cluster infrastructure; building solutions for deep, low-overhead system visibility; owning and evolving core observability platforms; building instrumentation libraries, SDKs, and eBPF auto-instrumentation; reducing mean time to detection/resolution through cross-signal correlation and AI-assisted diagnostics; driving fleet efficiency via continuous profiling and utilization insights; and partnering with Research, Inference, Product, and Infrastructure teams. You'll need hands-on experience building and operating large-scale observability infrastructure, deep end-to-end knowledge of observability signals (instrumentation through query), understanding of high-throughput telemetry tradeoffs, comfort digging below the application layer into kernel/network/hardware, excellent communication skills, and excitement for foundational infrastructure work. Preferred: 10+ years relevant experience, eBPF production experience, fleet-scale continuous profiling, kernel/syscall debugging, accelerator workload profiling, high-cardinality metrics systems, OpenTelemetry expertise, and interest in applying AI/LLMs to operational workflows.

Similar roles