SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Alpaca - Remote - Remote - posted 2026-08-05

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure serving hundreds of financial institutions across 40 countries. The company provides institutional-grade APIs for stocks, ETFs, options, crypto, fixed income, and 24/5 trading, supporting over 10 million brokerage accounts. Backed by $400M in funding from top-tier investors including Spark Capital, Tribe Capital, and Y Combinator, Alpaca operates a globally distributed team of 400+ engineers and brokerage professionals. As a Senior Site Reliability Engineer, you will help keep Alpaca's brokerage platform reliable, observable, and operable while the company scales. You'll work across cloud infrastructure, Kubernetes platforms, observability stacks, messaging layers, and data layers. The role emphasizes deep PostgreSQL expertise—the database sits on the trading-critical path, and you'll spend meaningful time leveling up the company's database reliability posture while remaining a well-rounded SRE. Key responsibilities include: operating production systems day-to-day (oncall, incident response, postmortems); defining and refining SLIs/SLOs and error budgets; strengthening observability across metrics, logs, traces, and alerting; shipping infrastructure through code in a GitOps workflow; managing PostgreSQL performance tuning, schema reviews, online migrations on large tables, HA/DR, and CDC pipelines; and mentoring engineers on reliability and database fundamentals. Required qualifications: 4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership; hands-on Kubernetes experience and GitOps infrastructure-as-code proficiency; solid PostgreSQL production knowledge (query plans, indexing, safe online migrations); cloud networking fundamentals (VPCs, routing, load balancing, DNS, TLS); modern observability stack comfort; Linux operator-level proficiency; incident response experience; working proficiency in Go or Python; and genuine interest in databases and PostgreSQL/DBA expertise growth. Nice-to-haves include deeper PostgreSQL experience (large OLTP clusters, online migrations, HA/DR ownership, connection pooling, CDC), typed SQL access layers in Go (pgx, gorm, sqlc), production messaging systems experience (RabbitMQ, Kafka, Redpanda), security/compliance in regulated environments (SOC 2, secrets management, audit logging), and familiarity with trading, brokerage, or fintech domains.

Similar roles