LMSYS Launches AutoEval for Rapid Model Evaluation
LMSYS (the Arena team) has officially introduced the AutoEval mechanism to its leaderboard, aiming to solve the time-consuming process of collecting real human votes. The system trains specialized Reward Models (RM) based on massive amounts of real user preference data, providing rapid first-day evaluation signals for newly released large models, which are then updated and verified once sufficient human votes are accumulated.
已确认
- 工作原理: AutoEval utilizes preference data from millions of real Arena users to train Reward Models, which automatically generate proxy votes to achieve second-level evaluations. The team has trained specialized reward models for different modalities (such as vision, image generation, and code).
- 性能表现: In text-to-image tasks, utilizing preference data exceeding a specific scale allowed its evaluation performance to surpass the baseline by 9 points.
- 有效性验证: The team compared AutoEval's early rankings with the final live leaderboard. Results showed that when the score gap between two models is significant, the accuracy of AutoEval's early rankings exceeds 90%.
为什么重要
AutoEval drastically shortens the model evaluation cycle, allowing new models to receive high-quality assessment signals calibrated with real data on their release day. This not only improves the leaderboard's update efficiency but also provides model developers with a reliable benchmark for comparing candidate checkpoints before the full leaderboard officially goes live.
2026-07-31 ~ 2026-07-31 · 5 related posts
Primary sources
- [source] Arena Validates AutoEval: Over 90% Accuracy in Ranking Models — arena · 2026-07-31
- [source] Arena Introduces AutoEval: Minute-Level Ratings via Reward Models — arena · 2026-07-31
- [source] Arena Details Multimodal Reward Models, Beating Baselines by 9 Points — arena · 2026-07-31
- LMSYS Launches AutoEval: Ranking Models via Reward Models in Hours — vista8 · 2026-07-31
1 near-duplicate retellings: arena