SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 120,000 - 150,000 / annual
Baselayer is rebuilding the identity and trust infrastructure for US businesses. The company has built the most comprehensive business graph in America by fusing public records, IRS data, sanctions lists, web signals, and fraud telemetry from 2,200+ financial institutions. With 98% match rates achieved in under two years (versus legacy bureaus at 60% over 50 years), Baselayer is trusted by over 20% of US financial institutions and is expanding into gig platforms, marketplaces, AI companies, and commerce infrastructure.
You'll join a small, technically deep team solving real-time entity resolution at scale—a graph AI problem, retrieval problem, and fraud-modeling problem combined. The team is at an inflection point: the graph is built, match rates are proven, and the hardest problems remain ahead: graph embeddings, fraud propagation models, sub-100ms latency traversal, and expanding identity beyond finance.
In this role, you'll build and maintain ETL/ELT pipelines that ingest and normalize data from dozens of sources. You'll develop data models and transformation layers using Dataflow, Spark, and Airflow that power fraud detection, KYB, and customer-facing APIs. You'll implement data quality checks, observability, and alerting to surface problems before customers see them. You'll tune pipelines and queries for performance, freshness, and cost in the cloud data warehouse. You'll collaborate with data scientists, ML engineers, and product teams to make clean, well-modeled data available for entity resolution and scoring. You'll help ensure pipelines meet security and regulatory standards for sensitive data (SOC 2, GDPR, KYC/KYB). You'll document your work and translate between technical and non-technical stakeholders.
This is a role for an early-career engineer who wants to be close to the action—feeding the models, not cleaning up after them. You'll write production code in your first weeks and own pipelines end to end, learning alongside senior data and ML engineers who invest in your growth. Ownership is real, velocity is real, and there's no layer of process between an idea and shipping it.
**REQUIREMENTS**
Minimum:
- 1+ years of experience in data engineering, working with Python, SQL, and cloud-native data platforms
- Experience building and maintaining ETL/ELT pipelines in a production environment
- Working knowledge of modern data stack tooling (e.g., Dataflow, Spark, Airflow, or equivalents)
- Hands-on experience with cloud data warehouses or lakes (e.g., BigQuery, Snowflake, or equivalents)
- Solid data modeling fundamentals and real care for data integrity and reliability
- Comfort with both structured and unstructured data, and a feel for what clean, scalable architecture looks like
What sets you apart:
- Curiosity about AI/ML infrastructure and a desire to be close to the models
- Experience with streaming or real-time data systems (e.g., Kafka, Pub/Sub)
- Exposure to KYC/KYB, fraud, risk, or underwriting data
- GCP experience (BigQuery, Cloud Run, Dataflow, Pub/Sub)
- Deep care for data quality and trust, building systems others can rely on
- Experience working without a playbook, taking direct feedback well and acting fast