SpatialClaw: Training-Free Code-as-Action Interface Boosts VLM Agentic Spatial Reasoning
CMHungSteven · x · 2026-09-29
A new arXiv paper from an NVIDIA internship team studies how the action interface shapes VLM spatial reasoning agents:
- Existing spatial agents use either single-pass code execution (locked into a strategy before seeing intermediate results) or structured tool-call interfaces (too rigid for freely composing operations).
- SpatialClaw is a training-free framework using code as the action interface: a stateful Python kernel preloaded with input frames and perception/geometry primitives, where the VLM writes one executable cell per step conditioned on all prior outputs — enabling flexible, open-ended 3D/4D spatial reasoning.
- The author invites agents/robotics/VLM eval researchers to reproduce the frontier-model numbers.
More from Research
- Amazon's GEB grounds entity biographies for long-video memory, hits 72% on EgoLifeQA — amazon · 2026-09-30
- NVIDIA's LongLive-Plug: Distill Once, Deploy Training-Free Across 54 Downstream Video Models — nvidia · 2026-09-30
- CrossBFM Distills a Shared Behavior Space Across Humanoid Robots in Under One GPU-Hour — Tan-Dzung Do · 2026-09-30
- AnisoWM: anisotropic representations improve planning in JEPA world models — SeoulNatlUniv · 2026-09-30
- Meituan's SAKI: Coupling-Routed Teacher Supervision Speeds Up On-Policy Distillation 4.22x — meituan · 2026-09-30
- Real2Gym Turns Videos Into Interactive Robot Gyms, Beating GPT-6 Direct Mode by 33% on Real Robots — Shanghai-AI-Laboratory · 2026-09-30