Spatial-Interactor teaches VLMs spatial reasoning through physical interaction
arankomatsuzaki · x · 2026-09-27
Researchers introduce Spatial-Interactor, a framework training VLMs to learn spatial reasoning and build spatial memory through interaction with the physical world, like an infant. Current spatial training focuses on static Q&A, giving little direct supervision for state transitions; interaction trajectories naturally link observation-action-observation pairs for direct supervision. Learning is organized into a three-level curriculum: passive world-state transitions, active self-state transitions, and long-horizon trajectories. The team also releases the LSI-108K dataset built from simulated and real interaction trajectories. Paper and code are public.
More from Embodied
- Claude Opus 5.5 Controls Robot Arm to Copy Michelangelo, Self-Corrects a Broken Line — 141_1337 · 2026-09-27
- Spider-legged welding robot: how staged protection tames arc-welding EMI — lukas_m_ziegler · 2026-09-27
- The humanoid industry is overrating legs: wheeled robots may win warehouses — VraserX · 2026-09-27
- 1X CEO targets 50,000 NEO humanoid robots shipped in 2027, 110,000/year capacity — AIFlow_ML · 2026-09-27
- Tesla Optimus scales to hundreds per week but hands remain a bottleneck — AIFlow_ML · 2026-09-27
- $800 litter-picking robot MOSS hits V0.3, open-source V0.4 already printing — AIFlow_ML · 2026-09-27