SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sardine is a leading agentic risk platform for fighting financial crime, offering integrated solutions that unify data across risk teams to stop fraud in real time, prevent AI-driven attacks, and automate fraud and AML operations. The company serves leading organizations including FIS, GoDaddy, Intuit, Edward Jones, ZoomInfo, and Checkout.com.
You will own the data and machine learning foundation that powers Sardine's compliance and onboarding decisions end-to-end. This is a high-impact, highly technical individual contributor role at the intersection of data engineering and ML engineering, requiring someone at the senior level to set technical direction for the next phase of growth.
Key responsibilities include:
- Own the data ingestion layer bringing device telemetry, transaction events, KYC/identity signals, and third-party enrichment into the platform. Design streaming pipelines (Pub/Sub, Apache Beam on Dataflow, Flink) and batch pipelines (Python, Airflow on Cloud Composer, Spark on Dataproc) that are correct, observable, and cost-efficient.
- Build and evolve the feature platform using Chronon feature definitions computed by Flink for streaming and Spark for batch, with aggregation windows from one hour to 300 days, served to the rules engine and models under sub-second latency budgets.
- Establish feature correctness as an engineering discipline through streaming-versus-batch reconciliation, recomputation tests, train/serve parity checks, and drift monitoring.
- Productionize fraud and identity ML models using Vertex AI and Kubeflow, gradient-boosted and tree-based models (XGBoost, LightGBM, CatBoost, scikit-learn), and build automated retraining, champion/challenger promotion, and rollback machinery.
- Engineer KYC, AML, and identity risk signals from document verification, sanctions/PEP screening, email and phone risk, synthetic identity indicators, and periodic customer due diligence.
- Integrate and harden new data sources, including 30+ third-party enrichment providers and the cross-client consortium network, owning failover behavior, timeout budgets, graceful degradation, and caching.
- Own the warehouse and modeling layer in BigQuery, including partitioning strategy, staging-to-mart architecture, training datasets, and migration from dbt to scheduled SQL and Python pipelines.
- Design entity resolution and graph data linking customers, devices, emails, phones, cards, bank accounts, and crypto addresses across clients.
- Ensure platform safety through field-level encryption, regional data residency enforcement, PII handling and deletion paths, and feature-level gating.
- Set technical direction and mentor engineers and data scientists through design docs, code reviews, and architectural decisions.
You will work directly with data scientists, backend engineers, and fraud analysts who use what you build. Sardine maintains a remote-first culture with hubs in the Bay Area, NYC, Austin, Toronto, and São Paulo, but you can work from anywhere in the US or Canada.