SpatialClaw: Training-Free Code-Action Agent Beats Prior by 11.2 Points on 20 Spatial Benchmarks
_akhaliq · x · 2026-09-30
NVIDIA researchers introduce SpatialClaw, a paper arguing that "code is the right action interface for spatial reasoning."
- Method: A VLM-backed agent writes Python in a persistent kernel, composing perception modules, inspecting intermediate results, and revising its strategy across steps. It is fully training-free, with no benchmark- or model-specific adaptation.
- Results: It beats a recent prior agent by +11.2 points across 20 benchmarks and improves consistently over six VLM backbones.
- Takeaway: Simply redesigning the agent's action interface around code execution yields large, general spatial-reasoning gains without any fine-tuning.
More from coding & agent
- Hugging Face datatrove 0.10.1 fixes empty-doc crashes and a silently ignored skip parameter — vanstriendaniel · 2026-09-30
- Arbor: open-source framework for AI agents doing autonomous long-horizon research — burkov · 2026-09-30
- Open-source coding agent Aster launches: self-hosted, any model, zero telemetry — saheedniyi_02 · 2026-09-30
- Gorgias built its 'company brain' Cortex with a 4-person team in 3-4 weeks — the year of groundwork mattered most — femke_plantinga · 2026-09-30
- Blocked Wi-Fi? Turn a Mac Mini into a Tailscale exit node to unblock Amp Code — iannuttall · 2026-09-30
- Grep Plus Claude: Using colgrep to Catch Dumb Code Patterns the Model Misses — antoine_chaffin · 2026-09-30