LMArena's post-training recipe lifts Flux2dev by 69 Elo on T2I leaderboard

lmarena-ai · hf · 2026-10-09

LMArena presents a post-training recipe for text-to-image models composing a Bradley-Terry preference reward with rubric-based rewards guarding against reward hacking, using a novel composition strategy instead of naive weighted averaging. RL-trained Flux2dev gains 69 Elo over base; post-trained Ideogram-4 reaches 1223.5 Elo, surpassing all open-source models on the Arena leaderboard. They also release a 1K training subset.

Original post →

More from Multimodal

Multimodal channel →