SCOPD Distillation Keeps 92% of VLM Performance at 10% Visual Tokens

uoft · hf · 2026-09-29

UofT researchers introduce SCOPD, a sparse-context on-policy self-distillation framework for efficient vision-language models.

Key insight: the representation-utilization gap

Method

Results

Original post →

More from Multimodal

Multimodal channel →