New translation benchmark spans 109 languages; LLM verifier pass rate jumps 7.2% to 89.8% with explicit rules
LChoshen · x · 2026-09-07
The "Last Translation Benchmark" paper, a collaboration spanning ETH Zürich, Edinburgh and dozens of institutions, builds 3,456 hard translation examples across 109 languages.
- Method: each example comes with explicit verification rules describing the specific error to avoid — wrong word sense, lost wordplay, gender resolution, cultural meaning, tone, or multimodal context — replacing vague quality scores with concrete failure checks.
- Key finding: when LLMs are given the human-written rules, verifier pass rate jumps from 7.2% to 89.8%, suggesting the bottleneck is often identifying the right translation constraint rather than satisfying it.
More from Research
- OR-Clarify benchmarks asking clarifying questions before optimization modeling — AIOR-Research · 2026-09-07
- How keyword spotting models power "Hey Siri": a detailed technical guide — cneuralnetwork · 2026-09-07
- Reward hacking stems from bad reward modeling and eval awareness, developer argues — secemp9 · 2026-09-07
- ECCV 2026 Oral Poppy: training-free polarization cues cut surface normal error by up to 26% — ssh4net · 2026-09-07
- AI OCR quietly corrupts protein sequences, and patent PDFs often contain the typos themselves — iskander · 2026-09-07
- NUS Study Finds LLMs Over-Edit Code; RL Boosts Minimal-Edit Fidelity Without Losing Accuracy — NationalUniversityofSingapore · 2026-09-07