SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Labrynth is an AI-powered platform company that helps organizations navigate regulatory complexity across heavily regulated industries like energy, compliance, and government. The company operates as a forward-deployed engineering organization with small, high-velocity teams embedded with clients.
You will join as a Data Engineer to build the data platform behind Labrynth's regulatory indices. The role spans three core modes: acquiring fragmented public data from government portals, APIs, PDFs, and protected sites; transforming inconsistent jurisdictional data through bronze-silver-gold medallion pipelines with idempotent ingestion and lineage tracking; and constructing transparent, auditable indices through statistical methodology including winsorization, percentile ranking, weighting, and sensitivity testing.
Key responsibilities include shipping resilient scrapers and ingestion flows against messy or adversarial data sources using HTTP/2 clients and browser automation; owning Postgres schema design and migrations across per-country and per-domain schemas; building and maintaining medallion transforms that are idempotent and content-hashed; implementing index methodology where the math is 100% test-covered; assessing data feasibility early and converting ambiguous requirements into executable plans; and taking indices end-to-end from sourcing through publication and refresh planning. You'll operate pipelines on Prefect orchestration (dispatching ECS Fargate tasks) with comprehensive observability.
The stack is deliberately modern: Python 3.14, uv, ruff, polars, Prefect 3, marimo, and AWS (ECS, S3, Terraform). You should have strong Python and SQL skills with production Postgres schema design experience; data pipeline expertise with lakehouse/medallion mindset; web scraping beyond basic requests (anti-bot evasion, browser automation); statistics literacy for index construction; modern Python tooling discipline (typing, linting, coverage gates); product discovery instincts to talk with partners and assess feasibility; and end-to-end ownership across data, backend, infrastructure, and product decisions in uncertain environments.
Nice-to-have skills include Prefect experience, AWS and Terraform proficiency, polars/pyarrow/marimo familiarity, LLM-in-pipeline experience, actuarial or quantitative research background, government open data experience, and comfort working alongside AI tooling in agent-forward repositories.