SlipstreamJobsFresh Startup & VC-Backed Jobs

Data & ML Engineer

DEFCON AI - Remote - Remote - posted 2026-08-06

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

DEFCON AI is building an AI-enabled decision-support system for resilient optimization of complex systems. As a Data & ML Engineer, you will own the data and model layer that powers this platform, operating in an accredited environment where every output must be traceable and defensible. The role spans four interconnected technical areas, and you'll indicate where your depth lies: **Data Modeling and Record Matching**: Design and maintain entity graphs with typed relationships. Implement probabilistic matching (blocking, candidate generation, pairwise scoring, clustering, threshold policy). Build deduplication and known-record suppression. Establish provenance so every node and edge traces to its source. Produce detailed interface and data-flow documentation. **Scoring and Calibration**: Develop relevance and priority models over large, imperfect record sets. Own calibration and threshold design—establishing what a score means, not just how it ranks. Design abstention policy to route uncertain and high-risk cases to humans. Perform feature engineering, establish baselines, and conduct error analysis accounting for asymmetric costs of false positives and negatives. **Retrieval and Generation**: Implement embeddings, vector storage, and retrieval across a provenance-tracked evidence base. Integrate language models through approved managed services and maintain open-weight alternatives. Design prompts and output schemas. Bind generated text to cited sources; treat "insufficient evidence" as a valid response. Own model packaging, serving, versioning, and rollback. **Pipelines and Source Handling**: Build secure ingestion, transformation, validation, and publishing across structured, semi-structured, and unstructured sources. Implement quality checks, schema validation, lineage capture, and audit logging. Detect source drift. Generate statistically representative synthetic data for development. You'll work to standards set by the Data Lead, document assumptions and caveats, instrument telemetry, maintain audit trails, and submit changes through a formal gate. The incoming data is predominantly low-signal, record matching is probabilistic, and false matches carry meaningful cost—making this a substantial technical challenge. You're not starting from scratch; an established platform exists with documented design decisions. Focus is on new capability: record matching, calibrated scoring, and grounded generation hardened for the target environment. This is fully remote with occasional travel (up to 25%) to HQ, customer sites, and vendor facilities.

Similar roles