SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Baseten is a Series F AI infrastructure company ($1.5B valuation) that powers mission-critical inference for leading AI companies including Cursor, Notion, Abridge, and Writer. The company provides applied AI research, flexible infrastructure, and developer tooling to help frontier AI companies bring cutting-edge models into production.
You will join the Observability team within the Infrastructure organization as an early member, directly shaping the observability experience for both internal and external customers. As Baseten scales across multiple cloud providers and diverse hardware, the volume and complexity of operational data is growing exponentially. Your team is responsible for building high-throughput ingest pipelines, cost-efficient storage solutions, and agentic diagnostic tools to enable detection, diagnosis, and resolution of issues in minutes rather than hours.
Key responsibilities include: designing and building scalable telemetry ingest and storage pipelines for metrics, logs, and traces across multi-cloud infrastructure; owning and evolving core observability platforms with migrations and architectural improvements; building instrumentation libraries, SDKs, and integrations that enable engineering teams to emit high-quality telemetry; driving alerting and SLO infrastructure for reliability monitoring; and partnering with Inference, Product, and Infrastructure teams to ensure observability solutions meet organizational needs.
You should have deep experience with at least one observability signal area (metrics, logging, tracing, or error analytics) and familiarity with others. Strong understanding of high-throughput data pipelines, columnar storage engines, and observability platforms like Prometheus, Grafana, ClickHouse, or OpenTelemetry is essential. Proficiency in Python, Rust, or Go is required. You'll need excellent communication skills, comfort working independently on ambiguous high-impact challenges, and interest in applying AI/LLMs to operational workflows such as automated root cause analysis and anomaly detection.