UMM-Reflection open-sourced: interleaved RL teaches unified models to self-correct images, GenEval 0.71→0.84

ziqi_huang_ · x · 2026-09-29

Ziqi Huang and collaborators open-sourced UMM-Reflection, a method that trains unified multimodal models to natively generate, reflect, and redraw via interleaved reinforcement learning: the model diagnoses errors in its own image, revises it, and iterates, with reflection text and image generation learned jointly inside one model.

How it works

Results: on BAGEL, GenEval improves 12.05 points over SFT (0.71 → 0.84), with transfer gains on unseen benchmarks: WISE (+10.97), T2I-CompBench++ (+4.63), and OneIG-Bench (+3.48).

Paper, code, models, and demo are all released.

Related event: UMM-Reflection: Interleaved RL Teaches Unified Multimodal Models to Self-Correct Images(3 posts)→

Original post →

More from Research

Research channel →