Arena Introduces AutoEval: Minute-Level Ratings via Reward Models
arena · x · 2026-07-31
Arena's official blog announced the introduction of AutoEval to its leaderboards to address the time-consuming nature of collecting real human votes.
- How it works: It trains a Reward Model (RM) on Arena's massive human preference dataset to automatically cast proxy votes. Combined with live human votes, this unlocks a model ranking in under an hour.
- Multimodal support: The mechanism extends beyond text to vision, image generation, and code. For instance, its text-to-image RM was trained on over 3 million preference pairs, achieving SOTA performance on the public MMRB2 benchmark.
- Dynamic updates: AutoEval scores provide a Day-1 signal for new model launches and are clearly labeled, later being updated and validated as sufficient human votes accumulate.
Related event: LMSYS Launches AutoEval for Second-Level LLM Evaluation(4 posts)→
More from Models
- Google Launches Gemini Robotics 2: Single Model Enables Multi-Robot Collaboration — DynamicWebPaige · 2026-07-31
- OpenAI Cuts Terra and Luna Model Prices, Luna Down by 80% — bindureddy · 2026-07-31
- Google Launches Gemini Robotics ER 2 Embodied Reasoning Model — rseroter · 2026-07-31
- Inkling-Small Released: 276B Parameter MoE Model Matches Original Performance — ziqiao_ma · 2026-07-31
- OpenAI Launches GPT-5.6: Luna Costs 80% Less, Terra 20% Less — kiki-le-koala · 2026-07-31
- OpenAI Kills Price Advantage of Chinese Open-Weight Models, Threatening Market Takeover — arrakis_ai · 2026-07-31