Paper finds systematic length bias in top translation metrics, skewing model rankings
shangbinfeng · x · 2026-10-04
- A paper to be presented at COLM (by @WendaXu2, @yilinjz et al., affiliated with Google AI/Google DeepMind) shows a systematic length bias across top-performing translation evaluation metrics and reward models.
- The authors argue these metrics may reward longer outputs rather than actual translation quality, fundamentally skewing model rankings—scores partly reflect sequence length, not quality.
More from Research
- Why LLMs suddenly excel at math and cybersecurity still has no predictive theory — burny_tech · 2026-10-04
- Cognitive-science abstractions transfer to AI: why Anthropic's interpretability looks psychological — sreejan_kumar · 2026-10-04
- Researcher gets 10 ICRA/RAL review requests in a week as robotics papers flood in — ChongZzZhang · 2026-10-04
- David Silver's 2024 slide resurfaces: the frontier recipe is just LLMs plus RL — akbirthko · 2026-10-04
- Christian Szegedy writes 'Quo Vadis, Mathematics?' on math's future in the AI era — ChrSzegedy · 2026-10-04
- COLM paper: thinking models amplify the wrong reasoning behaviors, 15,282 traces show — Jeande_d · 2026-10-04