Arena Details Multimodal Reward Models, Beating Baselines by 9 Points
arena · x · 2026-07-31
The LMSYS (Arena) team shared updates on their automated evaluation system (AutoEval), detailing the training of modality-specific reward models based on live platform preferences across vision, image generation, and code.
For Text-to-Image, their reward model was trained on over 3 million preference pairs. It achieves state-of-the-art performance among pointwise reward models on the public MMRB2 benchmark, outperforming the next-best baseline by more than 9 points.
Related event: LMSYS Launches AutoEval for Rapid Model Evaluation(5 posts)→
More from Multimodal
- Fish Audio Raises $52M Seed, Launches S2.1 Pro Voice Model — thisdudelikesAI · 2026-07-31
- KalpaLabs Launches Conversational Speech Model in Public Beta — ycombinator · 2026-07-31
- Testing Minimax H3: Cinematic Text VFX and Prompts — LudovicCreator · 2026-07-31
- FLUX 3 Preview Goes Live; Nous Research Launches Short Film Contest — NousResearch · 2026-07-31
- AI Startup Eigent Expands to US Hiring Core Team; MiniMax H3 Video Generation Stuns in Test — VoidAsuka · 2026-07-31
- Beginner Question: How to Handle Dataset Captions for LoRA Training? — Mean-Crab1827 · 2026-07-31