Last Translation Benchmark: 3,456 Crowdsourced Examples That Break SOTA MT

Researchers release the Last Translation Benchmark: 3,456 crowdsourced hard examples that defeat state-of-the-art MT models, with failures extending beyond figurative language and pointing to a need for better multilingual cultural reasoning.

2026-09-04 ~ 2026-09-04 · 3 related posts