Framing human eval as a bandit problem focuses annotation budget on top models

zouharvi · x · 2026-09-14

The arXiv paper "Dynamically Allocating Evaluation Effort for Model Ranking" formalizes multi-model human evaluation as best-arm identification with correlated arms in a multi-armed bandit setup, where pulling an arm means human-evaluating a model. By adaptively sampling based on intermediate rankings, annotation budget concentrates on the most competitive models. The authors prove the surprising A-optimality of a simple algorithm and show it improves discrimination between top-performing models, making evaluations faster, cheaper, and better aligned with large-scale competition goals. Implemented in Pearmut and used at scale in WMT26.

Original post →

More from Research

Research channel →