SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Braintrust is an agent observability platform used by teams at Notion, Stripe, Box, OpenAI, and Cloudflare to trace agents, identify issues, and run evaluations. The Platform Support team owns the technical front line for infrastructure, performance, and reliability of hybrid and self-hosted deployments.
In this role, you'll provide customer-facing support for Braintrust deployments across AWS, Azure, and GCP, from initial installation through steady-state operation. You'll debug real infrastructure problems including Kubernetes workloads, Terraform state, networking, VPC configuration, IAM, TLS, and cloud-provider-specific issues. Using logs, metrics, and traces, you'll diagnose performance and reliability issues in the backend—ingest throughput, query latency, database and object-store behavior—to identify root causes rather than symptoms.
You'll lead incident response for customer-impacting issues, triaging and communicating clearly while driving resolution. Rather than handing every problem to Engineering, you'll ship fixes by submitting PRs to backend services, Terraform modules, and deployment tooling. You'll build diagnostics, health checks, preflight validation, and self-service paths that enable customers to unblock themselves. You'll write and maintain runbooks and deployment documentation, turning hard-won answers into permanent solutions. You'll feed patterns back to Engineering and Product to prevent recurring failure modes, and participate in an on-call rotation for critical customer issues.
The ideal candidate has experience in a customer-facing technical role (Support Engineering, SRE, DevOps, Solutions Architecture, or Infrastructure Engineering) or backend/infrastructure engineering with genuine appetite for customer work. You need strong Kubernetes fundamentals—deploying, debugging, and scaling workloads, reading pod events and logs. Hands-on Terraform experience and depth in at least one major cloud (AWS preferred) are essential. You're comfortable in a backend codebase (Python, TypeScript, or Go) enough to reproduce bugs, trace them to source, and fix them. You're fluent with observability tooling and instinctively reach for data before opinion. Clear, calm, direct communication under pressure—especially with technical, blocked customers—is critical. You take problems personally and follow them until the customer is running again.
Bonus experience includes supporting self-hosted or on-prem enterprise software in regulated environments, multi-cloud experience (Azure or GCP alongside AWS), database and data-infrastructure depth (Postgres, ClickHouse, or similar analytical stores), observability or ML infrastructure experience, familiarity with LLM APIs and agent evaluation in production, or having built support/diagnostic tooling that measurably reduced ticket volume.