UMM-Reflection: interleaved RL teaches unified models to self-correct, GenEval 0.71→0.84
ziqi_huang_ · x · 2026-09-29
Open-sourced UMM-Reflection teaches unified multimodal models to generate, reflect, and redraw via interleaved RL. Sibling trajectories sharing one initial image enable group-relative advantage over reflection strategies, with a single trajectory-level advantage updating both reflection tokens and flow-based revisions — no verifier at inference. On BAGEL it lifts GenEval by 12.05 points over SFT (0.71→0.84), with gains transferring to WISE (+10.97), T2I-CompBench++ (+4.63) and OneIG (+3.48).
More from Research
- Meta's LSRM wins ECCV 2026 Best Paper Honorable Mention, scales 3D reconstruction with sparse attention — rsasaki0109 · 2026-09-30
- Amazon's GEB grounds entity biographies for long-video memory, hits 72% on EgoLifeQA — amazon · 2026-09-30
- NVIDIA's LongLive-Plug: Distill Once, Deploy Training-Free Across 54 Downstream Video Models — nvidia · 2026-09-30
- CrossBFM Distills a Shared Behavior Space Across Humanoid Robots in Under One GPU-Hour — Tan-Dzung Do · 2026-09-30
- AnisoWM: anisotropic representations improve planning in JEPA world models — SeoulNatlUniv · 2026-09-30
- Meituan's SAKI: Coupling-Routed Teacher Supervision Speeds Up On-Policy Distillation 4.22x — meituan · 2026-09-30