Spatial-Interactor teaches VLMs spatial reasoning through physical interaction

arankomatsuzaki · x · 2026-09-27

Researchers introduce Spatial-Interactor, a framework training VLMs to learn spatial reasoning and build spatial memory through interaction with the physical world, like an infant. Current spatial training focuses on static Q&A, giving little direct supervision for state transitions; interaction trajectories naturally link observation-action-observation pairs for direct supervision. Learning is organized into a three-level curriculum: passive world-state transitions, active self-state transitions, and long-horizon trajectories. The team also releases the LSI-108K dataset built from simulated and real interaction trajectories. Paper and code are public.

Original post →

More from Embodied

Embodied channel →