UniEvo-VL: on-policy self-distillation lifts Qwen-image GenEval from 0.747 to 0.808

yining_hong · x · 2026-10-02

UniEvo-VL explores multimodal self-improvement via on-policy self-distillation: as generation and understanding converge in one model, the model's own feedback on its outputs becomes a training signal. Built on Qwen-image-2512, it improves GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. Experiments with stronger external critics suggest further headroom, highlighting that producing feedback is only the beginning of the self-improvement loop.

Related event: Stanford's UniEvo-VL Lets Multimodal Models Teach Themselves(3 posts)→

Original post →

More from Multimodal

Multimodal channel →