Pearmut and two eval methods replace Appraise-era tooling at WMT, headed to EMNLP
zouharvi · x · 2026-09-14
Zouhar summarizes a year of work on evaluation methodology at scale, with three results appearing at EMNLP:
- Pearmut: a lightweight, reproducible human annotation platform replacing Appraise at WMT, supporting DA/ESA/MQM protocols and already used beyond MT, including speech and general text annotation;
- Dynamic annotation allocation: bandit-inspired, focusing annotation budget on better-performing candidates as annotation progresses;
- cESA protocol: judging multiple translations at once for faster, less noisy judgments.
All are designed around real needs of large-scale evaluations like WMT and IWSLT.
Related event: Pearmut Platform and cESA Protocol Head to EMNLP(3 posts)→
More from Research
- Swapping pretraining objective cuts entity-swap false-accepts from 46% to 5% with zero training — Reasonable_Royal_621 · 2026-09-14
- LeanDB: Theoric Labs builds a strongly typed Lean frontend for SQL databases — hargup13 · 2026-09-14
- DeepMind looks back on 15 years of AI game research, partners with EVE Online devs — arnicas · 2026-09-14
- ECCV paper demystifies video reasoning: diffusion steps hide parallel trajectories — ziqi_huang_ · 2026-09-14
- Should arXiv reject an AI-discovered cancer breakthrough that passes clinical trials? — IanArawjo · 2026-09-14
- cESA: judging multiple translations at once cuts human eval time and noise — zouharvi · 2026-09-14