LMArena's post-training recipe lifts Flux2dev by 69 Elo on T2I leaderboard
lmarena-ai · hf · 2026-10-09
LMArena presents a post-training recipe for text-to-image models composing a Bradley-Terry preference reward with rubric-based rewards guarding against reward hacking, using a novel composition strategy instead of naive weighted averaging. RL-trained Flux2dev gains 69 Elo over base; post-trained Ideogram-4 reaches 1223.5 Elo, surpassing all open-source models on the Arena leaderboard. They also release a 1K training subset.
More from Multimodal
- Seedance 2.5 AI video demo stuns users with crazy detail level — SimplyAnnisa · 2026-10-09
- Claude Motion ships with day-one HyperFrames Studio integration for video editing — sean_t_strong · 2026-10-09
- Whistle: an open 16.9 MB speech-to-text model that runs on CPU with 11 ms first token in 7 languages — solyarisoftware · 2026-10-09
- Open-source huashu-art-motion turns coding agents into art-animation studios, 2.6k stars — AlchainHust · 2026-10-09
- VOCALOID7 AI Megpoid voicebank based on Megumi Nakajima opens preorders with 8 voice types — CurieuxExplorer · 2026-10-09
- Iris-3B open-sourced: a 3B pixel-space generation and general vision learner — Total-Resort-3120 · 2026-10-09