SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure, serving hundreds of financial institutions across 40 countries with institutional-grade APIs for stocks, ETFs, options, crypto, and fixed income trading. The company processes hundreds of millions of events daily and is backed by $400M in funding from top-tier investors including Spark Capital, Tribe Capital, and Y Combinator.
As a Senior Data Engineer, you will design and build the next generation of Alpaca's Data Platform to support scaling to larger customers and new jurisdictions. You'll own core data infrastructure spanning batch and stream ingestion, transformation, and consumption layers for BI/Reporting, AI/agent interfaces, and external third-party integrations. The role encompasses the full data stack: financial transactions, customer data, API logs, system metrics, and augmented data that drive decision-making for internal and external stakeholders.
Key responsibilities include:
- Design, build, and evolve core data platform infrastructure including distributed query engines, orchestration, warehousing, and cataloging systems
- Own lakehouse infrastructure as code using Terraform and Ansible on Kubernetes
- Build and maintain low-latency streaming and CDC ingestion pipelines with batch paths landing in Apache Iceberg
- Develop and scale the BI landscape for performant, self-serve access to lakehouse data
- Enforce platform reliability through monitoring, alerting, on-call rotations, incident response, and SLAs
- Partner with DevOps, Analytics Engineering, and other teams to close infrastructure gaps
Required qualifications include 5+ years of Data Engineering experience with 2+ years building and operating scalable, low-latency data platforms handling >100M events/day. You must have strong hands-on experience with Kubernetes, Docker, Helm, and infrastructure-as-code tools (Terraform, Ansible, ArgoCD). Deep knowledge of distributed systems, open-source query engines (Trino/Presto), Apache Iceberg, streaming systems (Kafka, Redpanda, Debezium), orchestration (Airflow), and ELT tools (Airbyte) is essential. Proficiency in Python and SQL, plus GCP experience (GCS, Cloud Build, Cloud SQL, Dataproc) or equivalent cloud services, is required. The ideal candidate thrives in fast-paced startup environments and can adapt infrastructure to rapidly changing needs.
Nice-to-have skills include experience with semantic/metrics layers (Cube, dbt, Looker), transformation frameworks (dbt), reverse ETL (Hightouch), data catalog and lineage tools (OpenMetadata, Datahub), and data governance frameworks.