Framing human eval as a bandit problem focuses annotation budget on top models
zouharvi · x · 2026-09-14
The arXiv paper "Dynamically Allocating Evaluation Effort for Model Ranking" formalizes multi-model human evaluation as best-arm identification with correlated arms in a multi-armed bandit setup, where pulling an arm means human-evaluating a model. By adaptively sampling based on intermediate rankings, annotation budget concentrates on the most competitive models. The authors prove the surprising A-optimality of a simple algorithm and show it improves discrimination between top-performing models, making evaluations faster, cheaper, and better aligned with large-scale competition goals. Implemented in Pearmut and used at scale in WMT26.
More from Research
- Swapping pretraining objective cuts entity-swap false-accepts from 46% to 5% with zero training — Reasonable_Royal_621 · 2026-09-14
- LeanDB: Theoric Labs builds a strongly typed Lean frontend for SQL databases — hargup13 · 2026-09-14
- DeepMind looks back on 15 years of AI game research, partners with EVE Online devs — arnicas · 2026-09-14
- ECCV paper demystifies video reasoning: diffusion steps hide parallel trajectories — ziqi_huang_ · 2026-09-14
- Should arXiv reject an AI-discovered cancer breakthrough that passes clinical trials? — IanArawjo · 2026-09-14
- cESA: judging multiple translations at once cuts human eval time and noise — zouharvi · 2026-09-14