Last Translation Benchmark Seeks Hard-to-Translate Samples

According to @zouharvi, a new machine translation benchmark called the Last Translation Benchmark is being built around “hard-to-translate” examples. The project is notable because it explicitly targets two problems in current MT evaluation: many benchmarks are too easy or already saturated, while automatic metrics are brittle and human evaluation is costly and hard to reproduce.

Core design

As summarized by @zouharvi, the benchmark has two main design choices. First, it collects difficult translation inputs across three modalities: text, images, and audio. Second, each example must come with human-readable “verification rules,” intended to make future evaluations of model outputs more provable and stable. @zouharvi said the authors shared examples to show that the benchmark is aiming at translation difficulties closer to real-world use than standard test sets usually capture.

Progress and next steps

@zouharvi reported that the project has already accepted 500+ examples from 118 contributors, spanning many languages and dialects. The team plans to use these submissions for rolling dataset analyses and a paper. He also noted that the research platform will open to public contributors, with the first version planned for September 1.

2026-07-17 ~ 2026-07-17 · 7 related posts