OPD-V: Optimizing Visual Reasoning via Modality Balance as Privileged Info
Aniri · hf · 2026-08-06
Visual reasoning in MLLMs often suffers from 'Modality Imbalance,' where textual information dominates and visual features are ignored. This paper introduces OPD-V, a new visual On-Policy Self-Distillation (OPSD) paradigm.
- Finding: Constructs Positive/Negative Teachers to prove 'Modality Balance' can serve as privileged information.
- Trust Region: Defines a 'Modality-Balance Trust Region' using Positive Modality-Balance Logits Margins to select on-policy tokens for distillation.
- Results: Across 6 benchmarks, 4 MLLM backbones, and 5 post-training methods, OPD-V consistently improves reasoning while reducing training costs.
More from Models
- Developer Complains About Claude's Strict Safety Guardrails: A Buzzkill — daniel_mac8 · 2026-08-06
- Paul Graham Asks Why LLMs Excel at Math But Struggle at Writing: Expert Points to Verifiable Rewards — chris_j_paxton · 2026-08-06
- DeepSeek reportedly planning 'significant increase' for API pricing — AlyoshaV · 2026-08-06
- Elon Musk Announces Launch of Grok Imagine Image Generation — elonmusk · 2026-08-06
- Inkling-Small Ties for First in Audio Reasoning, Ranks High in Tool Use — simonguozirui · 2026-08-06
- Zuck Teases He Will 'Share More on Open Source' Soon — realmvp77 · 2026-08-06