SpatialCLI: Teaching VLMs to Reason With Spatial Tools, Then Internalize Them

Yang Zhou · hf · 2026-07-31

General VLMs can reason about overall tasks but often miss crucial visual details, while specialist vision models capture details but cannot translate them into task-level decisions. To bridge this gap, this paper introduces SpatialCLI.

The framework teaches VLMs to use spatial tools and progressively internalize these specialist perceptual capabilities in three stages:

Experiments show that on MindCube, SpatialCLI improves Qwen3-VL-8B's accuracy from 29.3% to 84.6% (with tools), surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.

Original post →

More from Research

Research channel →