VAD isolates visually supported corrections for multimodal distillation and beats direct teacher supervision on six benchmarks
Kangning Zhang · hf · 2026-08-04
What the paper does
The paper studies multimodal on-policy distillation (OPD), where a privileged-view teacher supervises student-generated trajectories.
Core idea: Visual Attribution Distillation (VAD)
Instead of treating teacher corrections as a single mixed signal, VAD tries to isolate the part that is actually attributable to visual evidence:
- It evaluates the same fixed teacher with and without the relevant evidence.
- The change in centered log-probabilities is used as a signed proxy for the evidence direction.
- The original correction is projected onto that proxy to separate an intervention-aligned component from an unexplained residual.
- Training then uses a student-anchored reconstructed target from the aligned component, with the privileged teacher only as a weak regularizer.
Results
Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms:
- direct privileged-view distillation
- visual-advantage weighting
Token-level and controlled-target analyses suggest the method better captures task-relevant visual corrections, especially when the evidence refutes a wrong answer.
More from Multimodal
- AI Video Ads Breakthrough: Seedance 2.5 Nails Makeup Application with Zero Hallucinations — FinanceYF5 · 2026-08-04
- MiniMax H3 ComfyUI test shows better video quality at higher resolution and steps — ayakitodev · 2026-08-04
- A surreal AI video titled “kurt cobain, tupac, take money 1080” goes viral — TMcFly · 2026-08-04
- Experimental Sol-Attn Triton support lands in ComfyUI, tested on 4090 and 5090 — bdsqlsz · 2026-08-04
- MiniMax H3 users report pixelated output even at its 1344×768 default resolution — PhilMcGraw · 2026-08-04
- A Stable Diffusion LoRA keeps producing increasingly absurd Haaland posters — CryptoChangeling69 · 2026-08-04