VCSD boosts Qwen3-VL on ViRL39K without external teachers or evidence
UMCP · hf · 2026-07-24
What it is
A new method called Visual Contrastive Self-Distillation (VCSD) removes the need for an external teacher or privileged visual signals in on-policy self-distillation for vision-language models.
How it works
- At each student-generated prefix, an EMA teacher is run twice under the same prompt/prefix:
- once with the original image
- once with a content-erased control image
- The token-level log-probability gap highlights tokens specifically supported by the visual content.
- VCSD sharpens the teacher’s original-image distribution within plausible support, then distills that distribution into the student.
Results
On ViRL39K, VCSD consistently beats matched OPSD across Qwen3-VL and Qwen3.5 variants.
- Qwen3-VL 2B: 62.27% → 67.04%
- Qwen3-VL 4B: 71.30% → 73.16%
- Qwen3-VL 8B: 72.51% → 76.26%
Why it matters
The method needs only the image and question, adds no inference-time cost, and avoids external teachers, answers, evidence signals, or reasoning traces.
More from Multimodal
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11
- Imperium Game Trailer Showcases AI Video Generation — keaslenyt · 2026-09-11
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11