SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure serving hundreds of financial institutions across 40 countries. The company provides institutional-grade APIs for stocks, ETFs, options, crypto, and fixed income trading, supporting over 10 million brokerage accounts. Backed by $400M in funding from top-tier investors including Spark Capital and Y Combinator, Alpaca operates a globally distributed team of 400+ engineers and traders.
As a Senior DevOps Engineer, you will design, build, and operate the infrastructure enabling Alpaca to scale globally while running trading-critical systems with confidence. You'll have autonomy to design solutions against clearly defined goals and shape those goals with the team.
Key responsibilities include: designing and evolving cloud architecture on GCP (networking, interconnects, IAM, high-availability topology) expressed entirely as Infrastructure-as-Code using Terraform with GitOps principles; building and owning CI/CD pipelines for IaC changes with Policy-as-Code guardrails, drift detection, and progressive rollout; advancing Platform-as-a-Product by building self-serve capabilities and paved paths for engineers; strengthening the observability stack (Prometheus, Thanos, Grafana, Loki, Tempo, Alertmanager); operating GKE clusters and infrastructure services including Helm-packaged workloads, message brokers (RabbitMQ, IBM MQ), and data stores; participating in Follow-The-Sun on-call model with alert triage, incident response, and blameless post-mortems; and embedding SRE practices (SLIs/SLOs, error budgets, capacity planning) into infrastructure operations.
Required qualifications: 5+ years in DevOps, Platform/Infrastructure, or SRE roles operating large-scale, high-availability, high-performance production systems; deep hands-on GCP cloud architecture design experience; strong Infrastructure-as-Code skills with Terraform across multiple environments; proven CI/CD pipeline building for IaC with automated plan/apply and Policy-as-Code; significant Kubernetes (ideally GKE) and Helm production experience; solid cloud and L3/L4-L7 networking fundamentals; hands-on experience with modern observability stacks.