See2Think: Do Multimodal Models Truly Leverage Intermediate Visual States for Reasoning?

Siyu Yan · hf · 2026-07-31

While multimodal LLMs increasingly use sketches and annotations as intermediate visual states during reasoning, it's unclear if they truly rely on them. Researchers introduced the See2Think framework to evaluate this, comprising two components:

Evaluations reveal that visual reasoning is highly model- and environment-dependent. Process analysis shows models usually select relevant visual operations, but faithful rendering remains the biggest bottleneck, and high feedback uptake doesn't guarantee accuracy gains. Under task-relevant corrupted feedback interventions, accuracy drops by over 10 percentage points, confirming behavioral dependence on visual states.

Original post →

More from Multimodal

Multimodal channel →