SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 145,000 - 250,000 / annual
OpenTeams, founded by Travis Oliphant (creator of NumPy and SciPy) and built by open-source ecosystem veterans, is hiring a Senior AI/ML Test and Evaluation Engineer to build and operate benchmarking and evaluation capabilities at the core of an AI platform serving enterprise and government customers.
In this role, you will design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows. You'll develop evaluation methodologies that combine automated metrics with structured human subject matter expert judgment, ensuring rigorous comparison of candidate capabilities against current operational baselines. Your work will be hands-on engineering on an open-source toolchain, with a focus on identifying what models get wrong—failure modes, limitations, and edge cases that matter in production.
Key responsibilities include: curating and recommending candidate benchmarks based on mission needs; producing defensible evaluation reports with documented limitations; defining common standards for benchmark expression and reporting; building lightweight expert-scoring workflows and measuring inter-reviewer agreement; supporting partner organizations integrating their capabilities with shared evaluation standards; and developing reference notebooks enabling analysts to run, interpret, and extend evaluations.
Your reports inform senior stakeholders and drive decisions about which capabilities are ready to field. This is a high-stakes role where evaluation rigor directly impacts mission outcomes. You'll participate in structured feedback sessions with end users, document technical approaches for government stakeholders, and contribute to platform evolution.
Required: 6+ years software or ML engineering experience, including 3+ years evaluating, benchmarking, or deploying ML models in production or applied research. Strong Python proficiency, hands-on experience with PyTorch and Hugging Face, experience developing evaluation harnesses and benchmark suites, expertise in evaluation metrics and statistical rigor, and proven ability to build repeatable, auditable evaluation pipelines with documented data provenance. Experience evaluating LLMs or agentic workflows using task-based, metric-based, or judgment-based scoring is essential. Strong written communication skills required to explain methodologies, results, and limitations to technical and nontechnical audiences.
U.S. citizenship required. Active security clearance strongly preferred; candidates without active clearance may be considered for unclassified work but must be eligible to obtain and maintain clearance. Travel up to 15% may be required, primarily to government facilities and between company locations. Position contingent upon contract award.