See2Think: Do Multimodal Models Truly Leverage Intermediate Visual States for Reasoning?
Siyu Yan · hf · 2026-07-31
While multimodal LLMs increasingly use sketches and annotations as intermediate visual states during reasoning, it's unclear if they truly rely on them. Researchers introduced the See2Think framework to evaluate this, comprising two components:
- See2ThinkBench: 1,200 open-ended, visually dependent problems across 12 task categories (2D structured, 3D scene, real-world reasoning).
- Visual Action-of-Thought (VAoT): Records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled settings.
Evaluations reveal that visual reasoning is highly model- and environment-dependent. Process analysis shows models usually select relevant visual operations, but faithful rendering remains the biggest bottleneck, and high feedback uptake doesn't guarantee accuracy gains. Under task-relevant corrupted feedback interventions, accuracy drops by over 10 percentage points, confirming behavioral dependence on visual states.
More from Multimodal
- DeepSeek-V4-Flash vs Gemini 3.6 Flash: Multimodal 3D Generation Showdown — teortaxesTex · 2026-07-31
- ByteDance Launches Seedance 2.5: 30s Generation, API Price Hiked Over 50% — 智东西 · 2026-07-31
- Fan Uses AI Video Tools to Create Live-Action Trailer for Expedition 33 — imjm · 2026-07-31
- Hailuo AI's New Video Model Excels at Slow Motion with Character References — LudovicCreator · 2026-07-31
- Testing MiniMax H3 Motion Control: Rivals Kling 3.0 — OneTrueTreasure · 2026-07-31
- Higgsfield Offers Unlimited Seedance 2.0 Trial, Teases 2.5 Release — eyishazyer · 2026-07-31