Stanford's UniEvo-VL Self-Distillation Lifts Qwen-image GenEval From 0.747 to 0.808

stanfordnlp · hf · 2026-10-01

Stanford NLP introduces UniEvo-VL, an on-policy self-distillation recipe where one multimodal model acts as both teacher (seeing its own critique as privileged info) and student, minimizing divergence between their diffusion distributions. Built on open-source Qwen-image-2512, it lifts GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. Stronger external critics (e.g., GPT5.6-Luna) suggest a higher self-evolution ceiling, and gains are uneven across tasks like text rendering.

Original post →

More from Multimodal

Multimodal channel →