SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Data Infrastructure Engineer

Faire Wholesale, Inc. - San Francisco, CA, United States - Hybrid - posted 2026-10-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 246,500 - 339,000 / annual

Faire is a technology wholesale platform connecting independent retailers globally with products from around the world. The Data Platform group supports all data-dependent teams across the company—Product Engineering, Data Science, Machine Learning, Analytics, Strategy, Finance, and Product—ensuring data is accurate, accessible, and queryable without requiring deep infrastructure knowledge. You will own the technical direction and implementation of Faire's data infrastructure, specifically the path data takes from production databases (CockroachDB and MySQL) into analytical stores (Snowflake and Databricks). The current system uses Fivetran, Kafka, Spark, and Airflow but has accumulated technical debt. You will design and build the next generation of this infrastructure, handling the hard technical problems personally while bringing the rest of the company along. Key responsibilities include: - Setting technical direction for data movement from production to analytical systems and owning the multi-year roadmap - Building the CDC and streaming ingestion layer: CockroachDB changefeeds and MySQL binlogs into Kafka, then into Iceberg tables on S3, handling ordering, deduplication, late data, schema changes, and backfills - Implementing data contracts and quality checks throughout the platform - Establishing ownership and SLAs on critical business datasets, integrating quality checks with tools like Anomalo and Monte Carlo - Operating Airflow and Fivetran at scale, with informed opinions on build-versus-buy decisions - Owning platform reliability: SLOs, on-call rotations, incident reviews, and preventing recurrence - Leading cross-team migrations to the new platform without disrupting dependent teams - Mentoring engineers to improve the team's data engineering practice This is a hands-on role where you'll be the person other teams consult when they need guidance on how data should move at Faire. Requirements: - Proven track record building and running data infrastructure at meaningful scale that other teams depended on, with experience setting technical direction - Deep expertise in change data capture and streaming ingestion from operational databases via Kafka, including hands-on experience with ordering, duplicates, snapshots, and schema evolution - Hands-on experience with lakehouse architectures on open table formats (Iceberg on S3 preferred); comfortable discussing partitioning, compaction, catalogs, and copy-on-write versus merge-on-read - Strong Spark skills and experience running Databricks and Snowflake against shared storage - Practical experience with data quality and observability, including data contracts, SLAs, and tools like Anomalo or Monte Carlo - Experience operating Airflow at scale and working with managed ingestion tools like Fivetran - Strong SQL and ability to model data for analyst and data scientist usability - Solid Python plus at least one of Kotlin, Java, Scala, or Go; experience shipping infrastructure on AWS with Terraform - Working understanding of data governance: access control, PII, retention and deletion, lineage, and audit - Track record leading cross-team data initiatives and migrations, mentoring senior engineers - Ability to explain technical tradeoffs to both leadership and junior engineers; skill in facilitating decisions among disagreeing parties - Ownership mindset for broken or unowned systems; willingness to be on-call for built systems - Experience in marketplace, e-commerce, or transaction-heavy businesses is a plus

Similar roles