UMM-Reflection: interleaved RL teaches unified models to self-correct, GenEval 0.71→0.84

ziqi_huang_ · x · 2026-09-29

Open-sourced UMM-Reflection teaches unified multimodal models to generate, reflect, and redraw via interleaved RL. Sibling trajectories sharing one initial image enable group-relative advantage over reflection strategies, with a single trajectory-level advantage updating both reflection tokens and flow-based revisions — no verifier at inference. On BAGEL it lifts GenEval by 12.05 points over SFT (0.71→0.84), with gains transferring to WISE (+10.97), T2I-CompBench++ (+4.63) and OneIG (+3.48).

Related event: UMM-Reflection: Interleaved RL Teaches Unified Multimodal Models to Self-Correct Images(3 posts)→

Original post →

More from Research

Research channel →