Last Translation Benchmark curates 3,456 examples that break state-of-the-art MT models
prajdabre · x · 2026-10-04
Researchers launched the Last Translation Benchmark (LTB), arguing that current MT benchmarks are saturated and metrics are unreliable or unscalable. Instead, LTB crowdsources inputs that demonstrably break state-of-the-art translation models, each paired with a robust automatic pass/fail verification rule.
- LTBv1 contains 3,456 examples spanning text, image, audio, and video inputs; the paper is on arXiv, with a leaderboard on Hugging Face
- A cited example: Indonesian uses "bapak" (father) as a respectful term for strangers, and MT mistranslates it literally, making everyone your dad — showing cultural context in lower-resource languages remains a blind spot
- Submissions are accepted on a rolling basis through end of 2026; contributors with 10 approved entries earn co-authorship on the live publication
- The post urges UK/European academics to pay attention, noting many still assume MT is solved because English-French works
More from Research
- Arthur Gretton to Talk on Gradient Flows on MMD at NYC Probabilistic Modeling Workshop — ArthurGretton · 2026-10-04
- Tartan IMU Challenge Draws 131 Teams, Top 10 to Present Solutions — GhaffariMaani · 2026-10-04
- How LLMs actually work: embeddings, inference dynamics and the autoregressive loop, explained — gerardsans · 2026-10-04
- Engineer pushes back on the Platonic Representation Hypothesis hype — gerardsans · 2026-10-04
- Daimon's Tactile World Model Threads Beads at IROS by Feel, Not Just Vision — CyberRobooo · 2026-10-04
- Distilling an LLM into two 287M GLiNER encoders for court-decision extraction — results fall just short of the teacher — SignificantZebra5883 · 2026-10-04