Arena's reward recipe lifts post-trained FLUX.2-dev by 69 Elo on live T2I leaderboard
arena · x · 2026-10-02
Arena details a post-training recipe for text-to-image models combining human preference rewards with rubric-based rewards: a Bradley-Terry reward model trained on 5.6M pairwise votes, a VLM-evaluated faithfulness reward from auto-generated checklists, constraint rewards covering explicit and implicit intent, and anti-reward-hacking rubrics targeting failures like garbled text and photorealism drift.
Results: post-trained FLUX.2-dev gains 69 Elo on Arena's live T2I leaderboard (1202); post-trained Ideogram 4 gains 20 Elo to 1224, surpassing all publicly listed open models. Offline ablations with Gemini 3.5 Flash as judge show complementary rewards push win rate to 64.2%, and weight-space ensembling of policies with/without the anti-reward-hacking objective raises it to 66.0%.
Related event: Arena unveils post-training method combining preference and rubric rewards(2 posts)→
More from Multimodal
- Full music video generated locally with ComfyUI and LTX 2.5 on 16GB VRAM — sokmech · 2026-10-02
- Eleven v4 character animation demo shown off in new video — buraktuyan · 2026-10-02
- Claude directs its own EDM music video 'The Good Ending' in Fable 5.1 demo — cheetoskull · 2026-10-02
- All-in-one Krea 2 Turbo ComfyUI workflow uses 2/4-step distilled LoRAs for speed — TimeTruth2490 · 2026-10-02
- No one on the team knew Blender — Claude ran the whole 3D film workflow — gen_ericai · 2026-10-02
- Meshy teases Meshy Edit: tweak specific parts of 3D models with a text prompt — rms80 · 2026-10-02