VCSD lets vision-language models self-distill from image-content contrast alone

burny_tech · x · 2026-07-27

Visual Contrastive Self-Distillation (VCSD) proposes a simpler way to do on-policy self-distillation for vision-language models by removing image content from the teacher signal without relying on external teacher models or privileged answers.

How it works

Results

Original post →

More from Research

Research channel →