Last Translation Benchmark: 3456 Crowdsourced Examples Break SOTA Translation Models

zouharvi · x · 2026-09-04

The Last Translation Benchmark paper is out: a massive crowdsourcing effort collected 3456 unique hard-to-translate examples that demonstrably break state-of-the-art translation models, enabling more reliable evaluation.

Related event: Last Translation Benchmark: 3,456 Crowdsourced Examples That Break SOTA MT(3 posts)→

Original post →

More from Research

Research channel →