SpatialCLI: Teaching VLMs to Reason With Spatial Tools, Then Internalize Them
Yang Zhou · hf · 2026-07-31
General VLMs can reason about overall tasks but often miss crucial visual details, while specialist vision models capture details but cannot translate them into task-level decisions. To bridge this gap, this paper introduces SpatialCLI.
The framework teaches VLMs to use spatial tools and progressively internalize these specialist perceptual capabilities in three stages:
- Call: Exposes specialist vision models as spatial tools.
- Learn: Uses Cold-Start SFT and agentic RL to improve tool use.
- Internalize: Verbalizes successful tool-use trajectories to internalize capabilities.
Experiments show that on MindCube, SpatialCLI improves Qwen3-VL-8B's accuracy from 29.3% to 84.6% (with tools), surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.
More from Research
- Meta Proposes OneShot Retrieval Framework, Deployed in Instagram — _reachsumit · 2026-07-31
- Google Proposes HA-MoE Architecture to Boost Cross-Content Ranking Fairness — _reachsumit · 2026-07-31
- Research: Enhancing Generative Recommendation with LLM-Derived Language Tags — _reachsumit · 2026-07-31
- Meta's ROCS Paradigm Boosts Recommendation Retrieval QPS Up to 3x — _reachsumit · 2026-07-31
- With Mandatory Reviewing, Low-Quality Reviews in AI Conferences Are No Longer Justifiable — Kwangryeol · 2026-07-31
- HiLaR: Optimizing LLM Recommendation Reasoning via Hierarchical RL — _reachsumit · 2026-07-31