cESA: judging multiple translations at once cuts human eval time and noise

zouharvi · x · 2026-09-14

The arXiv paper "Contrastive ESA: Human Evaluation of Multiple Translations at Once" proposes cESA (Contrastive Error Span Annotation): annotators view multiple translations of the same source side by side, mark major/minor error spans, and assign an absolute 0–100% score. Shared context across outputs yields more consistent and faster judgments than pointwise evaluation. Validated with a large-scale English→Japanese human evaluation of 12 models, cESA reduces annotation time and noise; unlike contrastive ranking, it produces absolute quality scores enabling simple non-parametric model rankings without post-hoc corrections. Implemented in Pearmut and used in WMT26.

Original post →

More from Research

Research channel →