Last Translation Benchmark Is Live: Contribute 10 Examples, Get Co-authorship
zouharvi · x · 2026-09-04
Last Translation Benchmark is a live paper+dataset still open for contributions. It tackles saturated benchmarks and unreliable metrics by collecting inputs (text, image, audio, video) that provably break state-of-the-art translation models, each with an automatic pass/fail verification rule. Contributors with 10 approved submissions earn co-authorship on the live publication. LTBv1 already has 3456 examples (90MB) on arXiv and a Hugging Face leaderboard, with rolling future releases.
More from Research
- POSTECH's PACE uses coordinated agents to surface hidden conflicts in user requests — POSTECH · 2026-09-04
- 1981 Sloman paper argued emotions are inevitable in machines juggling multiple motives — yeastsplainer · 2026-09-04
- Life Biosciences moves Sinclair's epigenetic reprogramming drug ER-100 into Phase 1 trial — Olivier__OG · 2026-09-04
- New Paper Tackles TCR Pairings and Binding Boundaries in Antigen Recognition Prediction — victorgreiff · 2026-09-04
- Terminal-Bench Science nears 70% saturation months after launch, dynamic evals needed — shyamalanadkat · 2026-09-04
- New Testable AGI Definition Puts GPT-4 at 27% and GPT-5 at 58% of the Way — davidpattersonx · 2026-09-04