UniEvo-VL trains image models to learn from their own mistakes, lifting GenEval 74.7% to 80.8%

mark_k · x · 2026-10-02

A new paper, UniEvo-VL, proposes multimodal self-improvement: the model generates an image, critiques what went wrong, and converts the feedback into corrective instructions. A teacher generator sees those instructions while the student sees only the original prompt; training transfers the correction benefit into the student's weights so improvements persist.

Results (Qwen-Image-2512 with Qwen-VL feedback):

Outcomes vary across tasks, with mixed results on text rendering. The author highlights the core idea: turning self-critique into lasting improvement — learning to avoid mistakes is far more powerful than spotting them.

Related event: Stanford NLP's UniEvo-VL Lets Multimodal Models Self-Improve via Self-Distillation(4 posts)→

Original post →

More from Multimodal

Multimodal channel →