Arena details post-training T2I models with 5M preference votes and VLM-built rubric rewards
arena · x · 2026-10-02
Arena's new blog explains its text-to-image post-training approach: a reward model trained on 5M pairwise human preference votes, plus rubrics auto-constructed by VLMs that check prompt-following, constraints, and known reward-hacking behaviors. Preference alone is insufficient—images can look great while missing objects, adding unrequested content, or drifting in style.
Results: post-trained FLUX.2-dev climbs 69 points to #2 on Arena's live T2I leaderboard; post-trained Ideogram 4 scores 1224, surpassing all publicly listed open-source models. The post covers rubric reward construction, the full offline eval setup, and before/after visuals of caught reward-hacking failures.
Related event: Arena unveils post-training method combining preference and rubric rewards(2 posts)→
More from Multimodal
- Runway-Generated Short Film 'Do You Ever Miss a Life You Never Lived?' Resonates — tlakomy · 2026-10-03
- ComfyUI v0.38 speeds up MiniMax H3 I2V by at least 16% on AMD, test shows — mwhjose · 2026-10-03
- LTX 2.3 is fast but often ignores prompts, and LoRA support is drying up — Fit_Satisfaction2953 · 2026-10-03
- Reddit user crafts anime-girl AI wingman video with Opus 5, Seedance and Nano Banana — EndlessMendless · 2026-10-03
- How real has AI video gotten? Reddit debates realism with demo footage — Solar-Guardian · 2026-10-03
- KAIST's World Observer gives world models movable panoramic eyes to track unseen regions — Scobleizer · 2026-10-03