Last Translation Benchmark Seeks Hard-to-Translate Samples
According to @zouharvi, a new machine translation benchmark called the Last Translation Benchmark is being built around “hard-to-translate” examples. The project is notable because it explicitly targets two problems in current MT evaluation: many benchmarks are too easy or already saturated, while automatic metrics are brittle and human evaluation is costly and hard to reproduce.
Core design
As summarized by @zouharvi, the benchmark has two main design choices. First, it collects difficult translation inputs across three modalities: text, images, and audio. Second, each example must come with human-readable “verification rules,” intended to make future evaluations of model outputs more provable and stable. @zouharvi said the authors shared examples to show that the benchmark is aiming at translation difficulties closer to real-world use than standard test sets usually capture.
Progress and next steps
@zouharvi reported that the project has already accepted 500+ examples from 118 contributors, spanning many languages and dialects. The team plans to use these submissions for rolling dataset analyses and a paper. He also noted that the research platform will open to public contributors, with the first version planned for September 1.
2026-07-17 ~ 2026-07-17 · 7 related posts
- New Translation Benchmark Gathers Hard Samples — zouharvi · 2026-07-17
- Why Machine Translation Benchmarks Need a Rework — zouharvi · 2026-07-17
- The Last Translation Benchmark Initiative — zouharvi · 2026-07-17
- [source] Evaluating Translations Using Validation Rules — zouharvi · 2026-07-17
- [source] Research Platform Opens to Public Contributions — zouharvi · 2026-07-17
- [source] Translation Benchmark Collects Over 500 Samples — zouharvi · 2026-07-17
- Machine Translation Still Needs Hard Examples — zouharvi · 2026-07-17