UMM-Reflection: interleaved RL teaches unified models self-correcting image generation
Yijia Fan · hf · 2026-09-29
UMM-Reflection applies reinforcement learning to complete reflection trajectories inside unified multimodal models that can both see and render images, enabling diagnose-render-diagnose loops.
- Sibling trajectories share one initial image so group-relative advantage compares reflection strategies; trajectory-level advantage updates both reflection tokens and flow-based revisions
- No external critic or inference-time verifier needed; SFT alone can't find high-success repair paths
- On BAGEL, improves GenEval by 12.05 points over SFT, transferring to unseen benchmarks: WISE (+10.97), OneIG-Bench (+3.48), T2I-CompBench++ (+4.63)
More from Multimodal
- Teaching Claude Opus to Annotate Speech Timing for ElevenLabs Voiceovers Proves Tricky — RileyRalmuto · 2026-09-30
- Alibaba DAMO unveils WorldAttention for efficient interactive video world models — Alibaba-DAMO-Academy · 2026-09-30
- NUS's MaLiang-Harness exposes the Program-to-Visual gap in code-driven image/video generation — NationalUniversityofSingapore · 2026-09-30
- Thinking Reward Model: rubric-first scoring sets open-source visual generation reward SOTA — Xuehai Bai · 2026-09-30
- FurE reconstructs editable 3D animal fur 10x faster without animal-fur datasets — Srinjay Sarkar · 2026-09-30
- 'Canadian Winter Chimera' AI-Generated Video Goes Viral on Reddit — omegaphallic · 2026-09-30