SlipstreamJobsFresh Startup & VC-Backed Jobs

AI Data Engineer, Data Platform

Collective - San Francisco, CA, USA - Hybrid - posted 2026-09-24

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Collective is a fintech platform empowering self-employed professionals with integrated business services—incorporation, accounting, bookkeeping, tax, and community. Backed by General Catalyst, Sound Ventures, QED Investors, and Google's Gradient Ventures, the company is featured in Forbes, Bloomberg, TechCrunch, and more. You will own and scale the data platform powering analytics, reporting, and AI across Collective. This is a hands-on, end-to-end role where you design, build, and maintain production data pipelines; model data into reliable, well-documented tables; and establish engineering standards that keep the platform trustworthy as the company grows. Key responsibilities: - Design and build scalable batch and event-driven pipelines ingesting data from application databases, SaaS tools, and external APIs into BigQuery using Fivetran, custom Python loaders, and orchestration tools. - Model data using dbt with dimensional and analytical designs, layered architecture (raw, staging, marts), clear grain, naming conventions, and comprehensive documentation. - Own data quality and reliability: implement testing, monitoring, alerting, and data contracts; define and meet freshness and accuracy SLAs; triage pipeline failures and data incidents to root cause. - Optimize warehouse performance and cost through query tuning, partitioning, clustering, and BigQuery spend management. - Establish engineering standards for version control, code review, CI/CD, and infrastructure-as-code; document systems and runbooks for team maintainability. - Govern and secure data with access controls, PII handling, and data retention practices appropriate for financial services; partner with Security and Legal on compliance. - Enable the business by partnering with product engineers on source schema design, and with analysts and stakeholders to translate business questions into reliable datasets, metric definitions, and self-serve reporting in Metabase. - Support AI and analytics use cases by maintaining the semantic layer, metric definitions, and documentation for LLM-based tools and internal agents to query the warehouse accurately. You will join the Data Engineering team within Engineering and work closely with product engineers, analysts, and business stakeholders across Operations, Finance, and Go-to-Market. Requirements: - 5+ years of professional experience in data engineering, analytics engineering, or closely related role, ideally at a B2B SaaS or fintech company. - Expert-level SQL and strong Python skills for building pipelines, transformations, and tooling; comfortable writing tested, production-grade code. - Hands-on production experience with a cloud data warehouse (BigQuery strongly preferred), dbt or equivalent transformation framework, managed ingestion tools (Fivetran or similar), and an orchestrator (Airflow, Dagster, Cloud Composer, or similar). - Deep understanding of dimensional modeling, layered warehouse architecture, and schema design with strong opinions on grain, naming, and consistency. - Experience implementing testing frameworks, lineage, monitoring, and alerting for data pipelines, and operating them in production including on-call. - Fluency with git-based workflows, code review, CI/CD, and infrastructure-as-code; treat data infrastructure as software. - Track record of taking ambiguous, high-impact problems and delivering reliable systems end-to-end with focus on outcomes. - Ability to explain technical trade-offs to non-technical stakeholders and drive alignment on data definitions across teams. Nice to have: - Experience with streaming or event data (Pub/Sub, Kafka, or similar) and product analytics tooling (Amplitude or similar). - Experience with Terraform and Google Cloud Platform infrastructure. - Experience with observability platforms such as Datadog. - Exposure to financial, accounting, tax, or payroll data and correctness requirements. - Experience building semantic layers or metric stores consumed by LLM-based tools, or supporting LLM evaluation programs. - AI-assisted development experience (Claude Code or similar).

Similar roles