SlipstreamJobsFresh Startup & VC-Backed Jobs

Model Test and Measurement Engineer

DEFCON AI - Remote - Remote - posted 2026-10-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 150,000 - 190,000 / annual

DEFCON AI is an insights company leveraging artificial intelligence, mathematical optimization, data analytics, and software engineering to optimize complex systems for resilience and better decision-making. As a Model Test and Measurement Engineer, you will own the independent evaluation framework that ensures AI and scoring systems perform as intended. You will build and maintain labeled ground truth datasets, design statistically sound audit and sampling methodologies, measure model performance across releases, and create evidence packages that support deployment decisions. Your role is to independently validate both scoring and generative AI capabilities, helping the team understand not only whether a model works, but how confidently its outputs can be trusted. This is a role with genuine influence. Your assessments will inform release decisions, drive improvement efforts, and provide objective evidence customers rely on when evaluating system performance. You will work closely with data scientists, AI engineers, and technical leadership while maintaining the independence needed to provide clear, defensible evaluations. Critically, you do not build the models you validate, and you do not approve thresholds or release authorization against your own evidence. Key Responsibilities: - Construct labeled ground truth for model evaluation - Design and run sampled audits - Gate releases against model version; maintain version inventory, evaluation records, and rollback triggers - Run drift and override review - Produce human-oversight and fairness/disparate-effect evidence - Maintain independence from build roles: produce evaluation evidence only - Measure workflow improvement: review time, throughput, backlog movement, override and rework rates - Design evaluation-phase QC: sample selection that does not mix populations, unit of review, and fair comparison when methods or searched sources differ This is a fully remote role with occasional travel to DEFCON AI headquarters, customer sites, and partner facilities as needed. Requirements: The posting does not explicitly state years of experience, specific technical certifications, or formal education requirements in the provided excerpt. However, the role requires expertise in statistical methodology, model evaluation frameworks, data labeling and ground truth construction, audit design, fairness and bias assessment in AI systems, and the ability to maintain independence while communicating technical findings to non-technical stakeholders.

Similar roles