SCOPD self-distillation recovers 92% of full-context VLM accuracy with 90% fewer visual tokens

CSProfKGD · x · 2026-10-01

Researchers observed that after aggressive visual token pruning, VLM Pass@1 drops but Pass@K recovers many failures—evidence survives, the model just fails to use it (the representation-utilization gap; 64 samples lift success from 53.2% to 79.6% on the same pruned tokens, vs 3-4% with no image). SCOPD closes the gap via on-policy self-distillation: a student reasons from pruned tokens while a full-context EMA teacher supervises the same trajectory, with SCOPD+ focusing supervision on vision-dependent tokens. On Qwen2.5-VL-7B with VisionZip, average score over 13 benchmarks rises from 86.37 to 92.43 at 10% tokens, +6.06 over pruning alone.

Related event: SCOPD recovers 92% of VLM performance after aggressive visual token pruning(2 posts)→

Original post →

More from Models

Models channel →