SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Binance, the world's largest cryptocurrency exchange by trading volume, is seeking a Senior Evaluation Algorithm Engineer to build and lead LLM evaluation capabilities for dialogue, financial trading, and other business scenarios.
In this role, you will design end-to-end evaluation plans that transform subjective model performance judgments into quantifiable, reproducible, and explainable conclusions. You will lead the design and construction of high-quality evaluation datasets, define evaluation dimensions and scenario coverage, establish annotation guidelines and quality control processes, and build benchmarks that authentically reflect business needs.
You will analyze model capability boundaries and failure modes based on evaluation results, produce actionable improvement recommendations, and collaborate with algorithm and product teams to drive model iteration. A key responsibility is driving the automation and scaling of evaluation workflows by building sustainable evaluation platforms and toolchains to support high-frequency, stable evaluation during rapid model iteration.
You will work cross-functionally with algorithm, product, and data teams to translate business and model objectives into clear evaluation standards, turning evaluation findings into concrete R&D directions and driving their implementation.
This is a high-impact role where evaluation serves as a critical component of the R&D loop, directly influencing the quality and direction of LLM development at a leading global blockchain ecosystem.
REQUIREMENTS:
- Master's degree or above in Computer Science, Artificial Intelligence, Mathematics, Statistics, or related fields, with solid algorithmic foundation and understanding of LLM principles, training, and fine-tuning processes
- Hands-on LLM evaluation experience at a large tech company, with participation in commercial deployment evaluation (not purely academic or offline benchmarking); familiar with full pipeline from evaluation data preparation and rubrics design to evaluation-driven R&D
- Familiarity with mainstream evaluation methods (human evaluation, model-based automatic evaluation/LLM-as-a-judge, metric computation) and their applicable boundaries; able to define appropriate evaluation dimensions for different business scenarios and write clear, actionable, discriminative rubrics
- Systematic control over evaluation data representativeness, annotation consistency, and result reliability, ensuring scientific and trustworthy evaluation conclusions
- Proficient in Python, with experience in evaluation workflow automation, benchmark construction, or evaluation platform development; able to independently handle data processing, evaluation script writing, and result analysis
- Strong business understanding and communication skills; able to translate evaluation findings into clear improvement directions and effectively drive cross-team collaboration
BONUS QUALIFICATIONS:
- Experience evaluating dialogue systems, AI Agents, or financial/trading LLMs
- Experience building high-quality AI training/evaluation data or data annotation systems
- Familiarity with RLHF, reward models, or preference data-related work