Arena Details Multimodal Reward Models, Beating Baselines by 9 Points

arena · x · 2026-07-31

The LMSYS (Arena) team shared updates on their automated evaluation system (AutoEval), detailing the training of modality-specific reward models based on live platform preferences across vision, image generation, and code.

For Text-to-Image, their reward model was trained on over 3 million preference pairs. It achieves state-of-the-art performance among pointwise reward models on the public MMRB2 benchmark, outperforming the next-best baseline by more than 9 points.

Related event: LMSYS Launches AutoEval for Rapid Model Evaluation(5 posts)→

Original post →

More from Multimodal

Multimodal channel →