Stanford's UniEvo-VL Self-Distillation Lifts Qwen-image GenEval From 0.747 to 0.808
stanfordnlp · hf · 2026-10-01
Stanford NLP introduces UniEvo-VL, an on-policy self-distillation recipe where one multimodal model acts as both teacher (seeing its own critique as privileged info) and student, minimizing divergence between their diffusion distributions. Built on open-source Qwen-image-2512, it lifts GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. Stronger external critics (e.g., GPT5.6-Luna) suggest a higher self-evolution ceiling, and gains are uneven across tasks like text rendering.
More from Multimodal
- Midjourney partial tone-reversal prompt turns morning photos into glowing darkroom stills — tisch_eins · 2026-10-01
- Aiden Bai shares a prompt that generates a 30-second game trailer with original music and sound effects — aidenybai · 2026-10-01
- Fan-made IronHide video created with MiniMax H3 — ChaoticBlast · 2026-10-01
- Google Arts & Culture launches nom nom: a Neural Cellular Automata world simulated by Gemini — zzznah · 2026-10-01
- Early user feedback: Suno v6 disappoints on music generation quality — hq4ai · 2026-10-01
- Looped DiT: 260M model beats 6.5x larger rival on T2I benchmarks with 4.9x less compute — arankomatsuzaki · 2026-10-01