UniEvo-VL: on-policy self-distillation lifts Qwen-image GenEval from 0.747 to 0.808
yining_hong · x · 2026-10-02
UniEvo-VL explores multimodal self-improvement via on-policy self-distillation: as generation and understanding converge in one model, the model's own feedback on its outputs becomes a training signal. Built on Qwen-image-2512, it improves GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. Experiments with stronger external critics suggest further headroom, highlighting that producing feedback is only the beginning of the self-improvement loop.
Related event: Stanford's UniEvo-VL Lets Multimodal Models Teach Themselves(3 posts)→
More from Multimodal
- Ben Affleck fine-tunes open video models on own-shot data; Netflix bought his AI firm for $587M — rohanpaul_ai · 2026-10-02
- ComfyUI node pack adds LLM prompt enhancer for Qwen, MiniMax and Ideogram workflows — Constant-Tower9662 · 2026-10-02
- MIT's InstructMesh fixes AI-generated 3D models for real fabrication — nordicinst · 2026-10-02
- Suno launches Speech beta, the first audio model generating voice with matching music — suno · 2026-10-02
- MiniMax H3 continuation clips break context: prompt generator has no memory of prior scene — SnooMacaroons1365 · 2026-10-02
- MiniMax H3 RefMods deep dive: video, audio, motion & style with full ComfyUI workflows — Citadel_Employee · 2026-10-02