SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ASAPP is hiring a Machine Learning Engineer to lead the technical development of an evaluation platform for agentic and LLM-based systems. This is a high-impact role combining hands-on engineering with technical leadership and mentorship.
You will own the technical roadmap and architecture for the evaluation platform, spanning offline benchmarking through online production monitoring. Key responsibilities include:
• Designing evaluation methodologies appropriate to different pipeline stages: golden/regression test sets, human-in-the-loop review workflows, LLM-as-judge approaches, and automated metrics for task success, safety, and hallucinations.
• Building the data infrastructure that evaluation depends on: annotation and labeling pipelines, dataset versioning, data quality checks, and self-service tooling for researchers and product teams to run and interpret experiments independently.
• Partnering closely with Research, Product, and Platform teams to productize experiments into robust AI solutions.
• Representing the eval platform to stakeholders, setting expectations for model/agent releases, and reporting on platform health and coverage.
• Staying current with advancements in ML, NLP, voice, and LLM systems, and contributing to technical discussions across teams.
• Mentoring and supporting other engineers through design reviews, feedback, and knowledge sharing.
This role requires deep, hands-on experience building and operating evaluation systems for modern ML/LLM/agentic systems—not just consuming existing tools. You will be accountable for the long-term health and architecture of a complex, data-intensive system.
REQUIREMENTS:
• Deep, hands-on experience building and operating evaluation systems for modern ML/LLM/agentic systems.
• Demonstrated experience leading technical direction of a project or small team: setting architecture, driving design reviews, and being accountable for system long-term health.
• Strong architectural skills with proven experience designing complex, data-intensive software systems.
• Production experience with Python, AWS, Kubernetes, and/or Docker.
• Experience designing data pipelines for ML evaluation: labeling/annotation workflows, dataset versioning, quality control, and reproducible benchmarking.
• Bachelor's Degree in CS or related field.
• Demonstrated technical mentorship of junior and mid-level engineers, driving adoption of best practices and architectural alignment.
• Desire to learn, teach, and collaborate closely with cross-functional peers.
NICE-TO-HAVE:
• Experience building and evaluating agentic systems at scale.
• Experience with voice/audio quality evaluations.
• Production experience with LLM-centric services (inference, orchestration, evaluation, monitoring).
• Familiarity with large-scale ML experimentation, benchmarking, or simulation frameworks.
• Experience with conversational/customer-support AI domains (containment rate, conversation quality, goal completion).
• Knowledge of techniques for optimizing model architectures for faster inference.
• Experience with AWS, CI/CD, Kafka, Athena.