SlipstreamJobsFresh Startup & VC-Backed Jobs

Platform Support Engineer

Braintrust - San Francisco, CA, USA - In-office - posted 2026-09-09

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Braintrust is an agent observability platform used by teams at Notion, Stripe, Box, OpenAI, and Cloudflare to trace agents, identify issues, and run evaluations. The Platform Support team owns the technical front line for infrastructure, performance, and reliability for customers running hybrid and self-hosted deployments. In this role, you'll be the technical point of contact for customers deploying Braintrust across AWS, Azure, and GCP. You'll debug real infrastructure problems including Kubernetes workloads, Terraform state, networking, VPC configuration, IAM, and TLS. You'll diagnose performance and reliability issues in the backend—ingest throughput, query latency, database behavior—using logs, metrics, and traces to identify root causes. Key responsibilities include leading incident response for customer-impacting issues, triaging and communicating clearly under pressure, and driving resolution. You'll ship fixes by submitting PRs to backend services, Terraform modules, and deployment tooling rather than handing problems off. You'll build diagnostics, health checks, preflight validation, and self-service tools that help customers unblock themselves. You'll write and maintain runbooks and deployment documentation, feed patterns back to Engineering and Product to prevent recurring failures, and participate in an on-call rotation for critical issues. You should have experience in a customer-facing technical role (Support Engineering, SRE, DevOps, Solutions Architecture, or Infrastructure Engineering) or backend/infra engineering with genuine appetite for customer work. Strong Kubernetes fundamentals are essential—you can deploy, debug, and scale workloads, and read pod events and logs. You need hands-on Terraform experience and depth in at least one major cloud (AWS strongly preferred). You should be comfortable in a backend codebase (Python, TypeScript, or Go) enough to reproduce bugs and fix them. Fluency with observability tooling and the instinct to reach for data before opinion are critical. Clear, calm, direct communication under pressure—especially with technical, blocked customers—is essential. You take problems personally and follow them until the customer is running again. Bonus experience includes supporting self-hosted or on-prem enterprise software in regulated environments, multi-cloud experience (Azure or GCP alongside AWS), database and data-infrastructure depth (Postgres, ClickHouse), observability or ML infrastructure experience, familiarity with LLM APIs and agent evaluation in production, and having built support or diagnostic tooling that measurably reduced ticket volume.

Similar roles