UniEvo-VL trains image models to learn from their own mistakes, lifting GenEval 74.7% to 80.8%
mark_k · x · 2026-10-02
A new paper, UniEvo-VL, proposes multimodal self-improvement: the model generates an image, critiques what went wrong, and converts the feedback into corrective instructions. A teacher generator sees those instructions while the student sees only the original prompt; training transfers the correction benefit into the student's weights so improvements persist.
Results (Qwen-Image-2512 with Qwen-VL feedback):
- GenEval rises from 74.7% to 80.8%
- Verifying proposed corrections before training pushes it to 81.8%
- A stronger external critic reaches 88.2%
Outcomes vary across tasks, with mixed results on text rendering. The author highlights the core idea: turning self-critique into lasting improvement — learning to avoid mistakes is far more powerful than spotting them.
More from Multimodal
- 4D Ride emerges as a new AI video benchmark after the Will Smith spaghetti era — Unlikely_Manager2495 · 2026-10-02
- Midjourney recipe: long exposure + low stylize yields cinematic 35mm film portraits — michaelrabone · 2026-10-02
- Prompt share: 3D Pixar-style mascot turnaround in three views — azed_ai · 2026-10-02
- Open-source ComfyUI extension loads Civitai workflows in one click, auto-downloads missing models — Zealousideal-Bee-300 · 2026-10-02
- Most People Use LLMs Shallowly in Filmmaking — Cinema-Grade AI Films Still Demand Real Craft — taherdhanera · 2026-10-02
- LTX 2.5 generates 60-second clips on a 32GB RTX 5090 as software optimization beats VRAM upgrades — OpenEffect3955 · 2026-10-02