SCOPD Distillation Keeps 92% of VLM Performance at 10% Visual Tokens
uoft · hf · 2026-09-29
UofT researchers introduce SCOPD, a sparse-context on-policy self-distillation framework for efficient vision-language models.
Key insight: the representation-utilization gap
- Conventional wisdom blames performance loss from visual-token pruning on irreversible loss of task-relevant information. A fixed-context Pass@K analysis shows otherwise: repeated sampling from the same pruned representation recovers many examples missed by greedy decoding—useful evidence remains, the model just uses it unreliably.
Method
- A student generates reasoning trajectories from pruned visual tokens while a privileged full-context teacher supervises the same on-policy prefixes
- No ground-truth responses, no architectural changes, no extra inference-time compute
- SCOPD+ uses a small visual-budget intervention to identify visually sensitive positions and selectively distill them
Results
- At 10% visual-token retention: Vanilla keeps 86.37% of unpruned performance across 13 benchmarks, SCOPD raises this to 90.49%, SCOPD+ to 92.43%
- Takeaway: efficient reasoning depends not just on which information survives pruning, but on how reliably the model learns to use it
More from Multimodal
- The more Suno you hear, the less you notice its flaws — context flooding in human-AI collaboration — voooooogel · 2026-09-29
- Testing Dreamina video generation with timestamped prompts and single-image reference — gen_ericai · 2026-09-29
- Testing Timestamped Prompts with Single Image Reference in dreaminacpp — gen_ericai · 2026-09-29
- Kling 4.0 Omni Reference demoed: multi-reference scene generation with instant language swaps — SarahAnnabels · 2026-09-29
- BananaStudio: Full-Resolution Nano Banana Image Studio Inside Gemini Canvas — Z3ROCOOL22 · 2026-09-29
- Suno Studio demo turns your voice into any instrument, starting with electric guitar — suno · 2026-09-29