Last Translation Benchmark seeks inputs that break SOTA translation models

zouharvi · x · 2026-10-08

The Last Translation Benchmark (LTB) is collecting submissions of inputs (text, image, audio, video) that provably break state-of-the-art translation models, each with a robust automatic pass/fail rule. Contributors get co-authorship on a living paper after 10 accepted submissions; LTBv1 has 3,456 examples with a HF leaderboard, and latest models added include Cohere North Small Translate, Mistral Large 4 and Qwen 3.8. Dataset and paper will roll out through end of 2026.

Related event: Last Translation Benchmark opens submissions, adds new models(2 posts)→

Original post →

More from Research

Research channel →