Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards
Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh
cs.CV, cs.AI
2026-10-02
Preference rewards from 5.6M Arena votes plus intent-gated rubrics lift FLUX.2-dev by 69 Elo to 1202 and Ideogram-4 to 1223.5, past every open model on the Sept 4, 2026 board.
Open-domain text-to-image models are already strong. Further post-training gains depend on a reward that covers both what people like and the hard constraints pairwise votes rarely mention. Preference models such as PickScore and HPSv2 collapse human comparisons into one score. Aesthetics come through. Whether every named object is present, whether the requested style survived, and whether stray text appeared do not.
Optimizing that score alone keeps raising the reward while image quality peaks and then drops. FLUX.2-dev shows this curve against both an Arena-trained reward and public PickScore. A faithfulness-only reward washes out detail and style. Adding the raw scores together lets the largest scale dominate, and the policy can learn to please the detector.
The optimizer is DiffusionNFT. It updates the policy on the forward diffusion process: the current network and the network that drew the sample define a better and a worse velocity, and a high reward pulls toward the better one. Images from the same prompt are weighted by rank inside that group. Base weights stay frozen. Only a LoRA adapter is trained.
The preference term is a Bradley-Terry model. Qwen3.6-27B reads a trailing classification token, and a linear head emits one logit. Pairs come from Arena between 16 April 2025 and 7 June 2026, about 5.6 million human votes across 100-plus text-to-image models. Color jitter and pixel noise reduce reliance on low-level image statistics.
Faithfulness asks whether requested content is present. gemini-3-pro-preview splits 10k training prompts into 182k yes/no questions; Qwen3.6-27B answers them, and the reward is the fraction of yeses. A constraint term penalizes content the prompt did not ask for, including explicit negatives and implicit ones, such as an extra object in a minimalist scene. gemini-3.1-pro-preview writes 26k of those questions. The same judge scores them.
Hack detectors act as vetoes. Garbled text is judged from the image alone. Unwanted photorealism, such as a storybook prompt drifting into a cinematic photo, is compared with a frozen-base image for the same prompt, in both orders. A violation removes any positive preference score on that sample and leaves a negative score unchanged. Clearing the detector adds nothing, so the detector does not become a new target.
Preference and faithfulness are always on. The constraint turns on only when an offline classifier marks the prompt strict. An Impressionist road may still grow flowers. Two authors blindly labeled 400 prompts: agreement is 77.8% on the three-way label and 89.5% on the strict-versus-rest cut the gate actually uses. Each axis is standardized inside the batch, then summed with weight 1.
Stage 1 trains preference, faithfulness, and the gated constraint. Stage 2 adds the veto and trains again. Stage 3 averages the two LoRA deltas. Prompts are real Arena requests with edits, one-line stubs, niche names, and non-English text removed: 10k for training, 1k for validation. FLUX.2-dev (32B) uses 48 groups of 24 rollouts, 10 steps, 512 resolution, and CFG 4.0. Ideogram-4 (9.3B) uses the same budget and a guidance weight of 7.0.
Offline evaluation follows MMRBv2. gemini-3.5-flash compares each checkpoint with the frozen base on 1k Arena validation prompts, and GPT-5.4 re-judges. The reported point is the best of steps 30, 60, 90, 120, and 150 under Gemini. Inference is at 1024. Gemini win rates against FLUX.2-dev:
| Method | Arena RM | PickScore |
| DiffusionNFT recipe | 0.290 | 0.290 |
| Faithfulness only | 0.366 | 0.366 |
| Preference only | 0.507 | 0.461 |
| Arena-T2I-Hard | 0.589 | 0.577 |
| Stage 1 | 0.642 | 0.581 |
| Stage 2 | 0.635 | 0.571 |
| Stage 3 | 0.660 | 0.626 |
Arena-T2I-Hard is preference plus faithfulness, without the gated constraint. The public DiffusionNFT schedule mixes PickScore, HPSv2, CLIPScore, GenEval, and OCR across stages and scores 0.290 here. GPT-5.4 puts Stage 3 at 0.572 and 0.623.
Stage 1 has 204 clear losses to the base, Stage 2 has 234, and 92 are shared. Stage 1's wins concentrate on composition (65%) and style (16%). Stage 2 wins more often on faithfulness (16% versus 9%) and clean text (9% versus 3%). After averaging, Stage 3 wins 0.560 of Stage 1's exclusive losses, 0.397 of Stage 2's, and 0.304 of prompts both parents lose. Averaging two checkpoints from the same run reaches 0.6435 with the Arena reward; averaging across recipes reaches 0.6598. With PickScore that contrast is 0.5610 versus 0.6256.
An additive hack reward scores 0.5260 with the Arena RM and 0.4194 with PickScore under Gemini. The one-sided veto scores 0.6346 and 0.5710. Without the intent gate those columns are 0.6046 and 0.5200; with the gate they are 0.6416 and 0.5810. Skipping the detector's applicability check leaves the Arena RM at 0.5591, against 0.6346 with the check. PickScore barely moves, from 0.5695 to 0.5710.
Downstream RL tracks how many votes the reward model saw. The same Stage 1 setup scores 0.537 at 10k pairs, 0.607 at 100k, 0.611 at 500k, and 0.642 at about 5.6 million.
Ideogram-4 gets only Stage 1, plus an OCR reward, because post-training hurt text rendering and no stable hack appeared. Under Gemini with the Arena reward, preference-only scores 0.266, Arena-T2I-Hard scores 0.401, and the gated constraint reaches 0.553. GPT-5.4 scores that Stage 1 run 0.623 on the same reward.
On the live Arena board, snapshot 4 September 2026, FLUX.2-dev moves from 1132.8±14 to 1190.3 at Stage 1 and 1202.1 at Stage 3 (listed as 1202±13), about 69 points over the base. Post-trained Ideogram-4 reaches 1223.5. Its base, also the previous best open model Ideogram-4.0-Quality, sits at 1204±4, 19.5 behind. Closed models stay higher: nano-banana-2 at 1261±5, seedream-5.0-pro at 1258±4, nano-banana-pro at 1232±5.
A 1k public subset, Arena-T2I-Training, retrains the setup with public PickScore. Gemini win rates are 0.566 and 0.615 for Stage 1 and Stage 3, against 0.581 and 0.626 on the full 10k prompts.
FLUX.2-dev and Ideogram-4, then the top open model, both beat their own bases under one setup with every reward weight left at 1. The part worth copying is the composition: standardize each axis, gate constraints by intent, veto hacks, and average LoRA deltas from different reward setups. Public PickScore in that same composition still reaches a Stage 3 win rate of 0.626. The 1k subset recovers most of the offline gain. It does not recover a reward model trained on 5.6 million Arena votes, or the 69-point live jump.
The paper tests two base models, skips tool-using multi-step workflows, does not have humans audit the decomposed questions, and reports one seed.
Offline scores are language models grading language models. Gemini writes the checklists and intent labels, Qwen3.6-27B scores them, and gemini-3.5-flash picks the checkpoint. The ablations have no live Elo. Ideogram-4's offline win rate is only 0.553, and its Elo gain is 19.5, much smaller than FLUX.2-dev's 69. Training rollouts are at 512; evaluation and the live board are at 1024.
Hack rules were written after an audit of 205 Stage 1 failures, and they cover garbled text and photo-style drift. The 0.290 DiffusionNFT baseline was built for a different prompt mix, so part of that gap is domain shift. The prose attaches gains of 7.1 and 4.9 points to the Stage 1 discussion; those gaps match Stage 3 against Arena-T2I-Hard (0.660 versus 0.589, and 0.626 versus 0.577). Stage 1 itself leads by about 5.3 and 0.4 points. Ranks are the 4 September 2026 snapshot.