SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure, serving hundreds of financial institutions across 40 countries with institutional-grade APIs for stocks, ETFs, options, crypto, and fixed income trading. The company processes hundreds of millions of events daily and is backed by $400M in funding from top-tier investors including Spark Capital, Tribe Capital, and Y Combinator.
As a Senior Data Engineer, you will design and build the next generation of Alpaca's Data Platform to support scaling to larger customers and new jurisdictions. You'll own core infrastructure spanning batch and stream ingestion, transformation, and consumption layers for BI/Reporting, AI/agent interfaces, and external third-party integrations. The role encompasses the full data stack: financial transactions, customer data, API logs, system metrics, and augmented data that drive decision-making for internal and external stakeholders.
Key responsibilities include: designing and evolving core data platform infrastructure (distributed query engines, orchestration, warehousing, cataloging); managing lakehouse infrastructure as code using Terraform and Ansible on Kubernetes; building and maintaining low-latency streaming and CDC ingestion pipelines with Iceberg; developing and scaling the BI landscape for self-serve data access; enforcing platform reliability through monitoring, alerting, on-call rotations, and SLAs; and partnering with DevOps and Analytics Engineering teams.
Required qualifications: 5+ years Data Engineering experience with 2+ years operating scalable, low-latency platforms handling 100M+ events/day; hands-on Kubernetes and cloud-native tooling (Docker, Helm); production IaC experience (Terraform, Ansible, ArgoCD); deep distributed systems knowledge with hands-on query engine experience (Trino/Presto); strong Apache Iceberg and object storage expertise; streaming/CDC systems experience (Kafka, Redpanda, Debezium); orchestration (Airflow) and ELT tools (Airbyte); Python and SQL proficiency; GCP data services experience or equivalent cloud background.
Nice-to-haves include semantic/metrics layers (Cube, dbt, Looker), dbt transformation frameworks, reverse ETL (Hightouch), data catalog/lineage tools (OpenMetadata, Datahub), and data governance experience. The team is 100% distributed and globally diverse, spanning USA, Canada, Japan, Hungary, Nigeria, Brazil, UK, and beyond.