Arena details post-training T2I models with 5M preference votes and VLM-built rubric rewards

arena · x · 2026-10-02

Arena's new blog explains its text-to-image post-training approach: a reward model trained on 5M pairwise human preference votes, plus rubrics auto-constructed by VLMs that check prompt-following, constraints, and known reward-hacking behaviors. Preference alone is insufficient—images can look great while missing objects, adding unrequested content, or drifting in style.

Results: post-trained FLUX.2-dev climbs 69 points to #2 on Arena's live T2I leaderboard; post-trained Ideogram 4 scores 1224, surpassing all publicly listed open-source models. The post covers rubric reward construction, the full offline eval setup, and before/after visuals of caught reward-hacking failures.

Related event: Arena unveils post-training method combining preference and rubric rewards(2 posts)→

Original post →

More from Multimodal

Multimodal channel →