LT-OPD On-Policy Self-Distillation Lifts 5% Visual-Token Retention to 82.3%
Junxian Li · hf · 2026-09-29
The paper proposes LT-OPD, an on-policy self-distillation framework for extreme visual token reduction in MLLMs.
- Method: the student rolls out responses with only a small fraction of visual tokens while a frozen full-token copy of the same model provides distributional supervision along student trajectories; a budget-level curriculum progressively decreases the token budget to stabilize training.
- Results: on nine benchmarks with Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, beating training-free, training-based, and RL baselines; gains transfer to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B.
- Efficiency: reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% with no extra inference overhead.
More from Research
- Pinductor uses LLM priors to learn POMDP world models with 330x fewer episodes than DreamerV3 — burny_tech · 2026-09-29
- MIT lab's PEM-UDE recovers interpretable chaos equations from noisy data — burny_tech · 2026-09-29
- SIGReg regularizer in LeJEPA and LeWM reveals hidden contrastive pairwise repulsion in JEPAs — burny_tech · 2026-09-29
- New article argues AI models are likely more capable than they appear — zetalyrae · 2026-09-29
- NanoGPT embedding table gains validated by earlier ngrammer paper, both using AdaGrad — _arohan_ · 2026-09-29
- AI-Generated Proofs Deserve Public Posting, Argues Researcher, Readers Can Just Ask AI to Rewrite Them — aran_nayebi · 2026-09-29