New Translation Benchmark Gathers Hard Samples

zouharvi · x · 2026-07-17

They point out two major issues with existing machine translation benchmarks: - Many tasks are too simple or saturated to drive further progress in the field. - Automated metrics are often fragile, while human evaluation is expensive and hard to reproduce. To address this, they are pushing a new translation benchmark with core strategies: - Collecting "hard-to-translate" inputs across text, image, and audio formats. - Requiring human-readable "validation rules" so future translation results can be evaluated more reliably. The author also calls for submissions of such samples to help build the subsequent paper and rolling dataset releases.

Related event: Last Translation Benchmark Seeks Hard-to-Translate Samples(7 posts)→

Original post →

More from Research

Research channel →