Last Translation Benchmark curates 3,456 examples that break state-of-the-art MT models

prajdabre · x · 2026-10-04

Researchers launched the Last Translation Benchmark (LTB), arguing that current MT benchmarks are saturated and metrics are unreliable or unscalable. Instead, LTB crowdsources inputs that demonstrably break state-of-the-art translation models, each paired with a robust automatic pass/fail verification rule.

Original post →

More from Research

Research channel →