Last Translation Benchmark seeks inputs that break SOTA translation models
zouharvi · x · 2026-10-08
The Last Translation Benchmark (LTB) is collecting submissions of inputs (text, image, audio, video) that provably break state-of-the-art translation models, each with a robust automatic pass/fail rule. Contributors get co-authorship on a living paper after 10 accepted submissions; LTBv1 has 3,456 examples with a HF leaderboard, and latest models added include Cohere North Small Translate, Mistral Large 4 and Qwen 3.8. Dataset and paper will roll out through end of 2026.
Related event: Last Translation Benchmark opens submissions, adds new models(2 posts)→
More from Research
- LeCun: automated formal proofs open a new era for math; critics say question taste can't automate — AryHHAry · 2026-10-08
- Terence Tao shares AHM statement: OpenAI's 700-file math release is "a demonstration of power, not scholarship" — burny_tech · 2026-10-08
- RSIGym treats recursive self-improvement as systems engineering; Opus 5 leads at 0.4809 — teortaxesTex · 2026-10-08
- DIRU: dendrite-inspired recurrent units improve chaotic forecasting and neonatal epilepsy detection — Dr_Alex_Crimi · 2026-10-08
- Formalization agents auto-generate errata for papers, says mathematician Armstrong — burny_tech · 2026-10-08
- New arXiv Paper Shows Lean Formalization Agents Deviate from Original Proofs — burny_tech · 2026-10-08