Last Translation Benchmark: 3456 Crowdsourced Examples Break SOTA Translation Models
zouharvi · x · 2026-09-04
The Last Translation Benchmark paper is out: a massive crowdsourcing effort collected 3456 unique hard-to-translate examples that demonstrably break state-of-the-art translation models, enabling more reliable evaluation.
- Motivation: existing MT benchmarks are saturated and metrics are unreliable or unscalable, so the project curates inputs (text, image, audio, video) that provably fail modern models
- LTBv1 (3456 examples, 90MB) is live on arXiv with a Hugging Face leaderboard
- It's a live paper+dataset: contributors submit examples with pass/fail verification rules; 10 approved submissions earn co-authorship, with rolling dataset and paper updates on GitHub
Related event: Last Translation Benchmark: 3,456 Crowdsourced Examples That Break SOTA MT(3 posts)→
More from Research
- POSTECH's PACE uses coordinated agents to surface hidden conflicts in user requests — POSTECH · 2026-09-04
- 1981 Sloman paper argued emotions are inevitable in machines juggling multiple motives — yeastsplainer · 2026-09-04
- Life Biosciences moves Sinclair's epigenetic reprogramming drug ER-100 into Phase 1 trial — Olivier__OG · 2026-09-04
- New Paper Tackles TCR Pairings and Binding Boundaries in Antigen Recognition Prediction — victorgreiff · 2026-09-04
- Terminal-Bench Science nears 70% saturation months after launch, dynamic evals needed — shyamalanadkat · 2026-09-04
- New Testable AGI Definition Puts GPT-4 at 27% and GPT-5 at 58% of the Way — davidpattersonx · 2026-09-04