Arena's reward recipe lifts post-trained FLUX.2-dev by 69 Elo on live T2I leaderboard

arena · x · 2026-10-02

Arena details a post-training recipe for text-to-image models combining human preference rewards with rubric-based rewards: a Bradley-Terry reward model trained on 5.6M pairwise votes, a VLM-evaluated faithfulness reward from auto-generated checklists, constraint rewards covering explicit and implicit intent, and anti-reward-hacking rubrics targeting failures like garbled text and photorealism drift.

Results: post-trained FLUX.2-dev gains 69 Elo on Arena's live T2I leaderboard (1202); post-trained Ideogram 4 gains 20 Elo to 1224, surpassing all publicly listed open models. Offline ablations with Gemini 3.5 Flash as judge show complementary rewards push win rate to 64.2%, and weight-space ensembling of policies with/without the anti-reward-hacking objective raises it to 66.0%.

Related event: Arena unveils post-training method combining preference and rubric rewards(2 posts)→

Original post →

More from Multimodal

Multimodal channel →