cESA: judging multiple translations at once cuts human eval time and noise
zouharvi · x · 2026-09-14
The arXiv paper "Contrastive ESA: Human Evaluation of Multiple Translations at Once" proposes cESA (Contrastive Error Span Annotation): annotators view multiple translations of the same source side by side, mark major/minor error spans, and assign an absolute 0–100% score. Shared context across outputs yields more consistent and faster judgments than pointwise evaluation. Validated with a large-scale English→Japanese human evaluation of 12 models, cESA reduces annotation time and noise; unlike contrastive ranking, it produces absolute quality scores enabling simple non-parametric model rankings without post-hoc corrections. Implemented in Pearmut and used in WMT26.
More from Research
- Swapping pretraining objective cuts entity-swap false-accepts from 46% to 5% with zero training — Reasonable_Royal_621 · 2026-09-14
- LeanDB: Theoric Labs builds a strongly typed Lean frontend for SQL databases — hargup13 · 2026-09-14
- DeepMind looks back on 15 years of AI game research, partners with EVE Online devs — arnicas · 2026-09-14
- ECCV paper demystifies video reasoning: diffusion steps hide parallel trajectories — ziqi_huang_ · 2026-09-14
- Should arXiv reject an AI-discovered cancer breakthrough that passes clinical trials? — IanArawjo · 2026-09-14
- Framing human eval as a bandit problem focuses annotation budget on top models — zouharvi · 2026-09-14