Why MLLMs Ignore Images: New Research Localizes the Vision-vs-Prior Bottleneck
Jiaang Li · hf · 2026-08-04
The paper investigates why Multimodal Large Language Models (MLLMs) sometimes ignore visual evidence in favor of pretrained language priors.
- Diagnostics & Benchmark: Introduces two diagnostic paradigms using image reconstruction to probe visual availability, and releases WhatIfVis, a benchmark spanning 5 dimensions (spatial-temporal, color, count, size, weight) to measure multimodal context sensitivity.
- Key Findings: Coarse-grained visual evidence is actually well-encoded; the bottleneck lies in post-perceptual utilization. Vanilla models show unstable sensitivity even with explicit instructions, while Supervised Fine-Tuning (SFT) significantly improves controllability.
- Steering Mechanism: Activation patching localizes the trade-off to architecture-specific depths. The vision-versus-prior trade-off is controllable along a learned steering vector, improving reliability even without intent instructions.
More from Research
- Profluent's New CRISPR Approach Expands Targetable Mutations by 10X — nathanbenaich · 2026-08-04
- Essay argues the 'Stochastic Parrot' concept is wrong and harms AI ethics — _FelixSimon_ · 2026-08-04
- MIT Proposes Reusable Failure Analysis Framework for Multimodal Clinical AI — MIT · 2026-08-04
- Eric Horvitz Proposes Decision-Analytic Steering for High-Stakes LM Decisions — erichorvitz · 2026-08-04
- Nature Medicine Study: LLMs Excel at Triage Discrimination But Lack Calibration — erichorvitz · 2026-08-04
- OpenBMB Launches Dual-Agent System to Automate Supercomputing Acceleration — aigclink · 2026-08-04