UMM-Reflection open-sourced: interleaved RL teaches unified models to self-correct images, GenEval 0.71→0.84
ziqi_huang_ · x · 2026-09-29
Ziqi Huang and collaborators open-sourced UMM-Reflection, a method that trains unified multimodal models to natively generate, reflect, and redraw via interleaved reinforcement learning: the model diagnoses errors in its own image, revises it, and iterates, with reflection text and image generation learned jointly inside one model.
How it works
- RL over complete reflection trajectories: sibling trajectories share one initial image, and a group-relative advantage compares reflection strategies; a trajectory-level advantage updates both reflection tokens and flow-based revisions, avoiding combinatorial blow-up of per-round credit assignment.
- Credit flows across rounds and across both roles of the same model; no external verifier needed at inference.
- SFT on reflection trajectories gives only a cold start; RL is what finds high-success repair paths.
Results: on BAGEL, GenEval improves 12.05 points over SFT (0.71 → 0.84), with transfer gains on unseen benchmarks: WISE (+10.97), T2I-CompBench++ (+4.63), and OneIG-Bench (+3.48).
Paper, code, models, and demo are all released.
More from Research
- EasyPPO: just fix the critic — stable PPO for LLM post-training with zero training collapse — teortaxesTex · 2026-09-30
- Looped MoE Tuning: 2x Experts, 0.5x Looped Layers, 2x Loops, Attention Untied — burny_tech · 2026-09-30
- NeuralFieldManifold accepted at NeurIPS 2026, extending neural manifolds to LFP/EEG — burny_tech · 2026-09-30
- GPU dominance in deep learning is largely accidental: the case for custom chiplet inference hardware — ai · 2026-09-30
- Study: Positive steering vectors skew LLM revealed preferences toward 'tainted' choices — repligate · 2026-09-30
- Models preferentially remove negative steering vectors without knowing what they do — repligate · 2026-09-30