VCSD boosts Qwen3-VL on ViRL39K without external teachers or evidence

UMCP · hf · 2026-07-24

What it is

A new method called Visual Contrastive Self-Distillation (VCSD) removes the need for an external teacher or privileged visual signals in on-policy self-distillation for vision-language models.

How it works

Results

On ViRL39K, VCSD consistently beats matched OPSD across Qwen3-VL and Qwen3.5 variants.

Why it matters

The method needs only the image and question, adds no inference-time cost, and avoids external teachers, answers, evidence signals, or reasoning traces.

Original post →

More from Multimodal

Multimodal channel →