SCOPD self-distillation recovers 92% of full-context VLM accuracy with 90% fewer visual tokens
CSProfKGD · x · 2026-10-01
Researchers observed that after aggressive visual token pruning, VLM Pass@1 drops but Pass@K recovers many failures—evidence survives, the model just fails to use it (the representation-utilization gap; 64 samples lift success from 53.2% to 79.6% on the same pruned tokens, vs 3-4% with no image). SCOPD closes the gap via on-policy self-distillation: a student reasons from pruned tokens while a full-context EMA teacher supervises the same trajectory, with SCOPD+ focusing supervision on vision-dependent tokens. On Qwen2.5-VL-7B with VisionZip, average score over 13 benchmarks rises from 86.37 to 92.43 at 10% tokens, +6.06 over pruning alone.
Related event: SCOPD recovers 92% of VLM performance after aggressive visual token pruning(2 posts)→
More from Models
- Dwarfstar's Bespoke Quants Run Qwen Fast on a 96GB M3 Ultra — TheRealJesus2 · 2026-10-01
- Codex lead says usage on the primary dot is virtually unlimited, new limits needed — taherdhanera · 2026-10-01
- Claude Opus 5.5 Tops Epoch Capabilities Index at 167, Edging Out GPT-6 Astra — scaling01 · 2026-10-01
- DeepSeek V4.1-Flash squeezes KV cache to 890 bytes per token, cuts persistent cache 8x — jbhuang0604 · 2026-10-01
- Fable 5.1 vibe check: stronger coding, half the tokens of Opus 5, and it finally talks like a person — every · 2026-10-01
- All Three Labs Shipped Frontier Models in 10 Days — How Do You Pick One for Production? — Sur_AI_guy · 2026-10-01