SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Baselayer is rebuilding the identity and verification infrastructure for US businesses. The company has built the most comprehensive business graph in America by fusing public records, IRS data, sanctions lists, web signals, and fraud telemetry from 2,200+ financial institutions, achieving 98% match rates in under two years. The platform is trusted by over 20% of US financial institutions and is expanding beyond finance into gig platforms, marketplaces, and AI companies.
As a Data Engineer, you'll build and maintain the ETL/ELT pipelines that power this identity layer. You'll work on ingesting and normalizing data from dozens of sources, developing data models and transformation layers using tools like Dataflow, Spark, and Airflow, and implementing data quality checks and observability tooling. Your responsibilities include tuning pipelines for performance and cost in the cloud data warehouse, collaborating with data scientists and ML engineers to support entity resolution and scoring, and ensuring compliance with security and regulatory standards (SOC 2, GDPR, KYC/KYB).
This is an early-career role positioned close to the action—you'll write production code in your first weeks, own pipelines end-to-end, and learn from senior data and ML engineers. The team is small, ownership is real, and there's minimal process between ideas and shipping. The technical problems are substantial: real-time entity resolution at scale, graph embeddings, fraud propagation modeling, and sub-100ms latency traversal.
Minimum qualifications include 1+ years of data engineering experience with Python, SQL, and cloud-native platforms; hands-on ETL/ELT pipeline work in production; familiarity with modern data stack tools; and experience with cloud data warehouses like BigQuery or Snowflake. Preferred experience includes streaming/real-time systems (Kafka, Pub/Sub), KYC/KYB or fraud domain knowledge, GCP expertise, and a demonstrated commitment to data quality and reliability.