SpatialCORE turns grounding confidence into a learning signal for spatial reasoning, hitting SOTA
Rafi Ibn Sultan · hf · 2026-10-01
Spatial reasoning remains a weakness of large vision-language models; existing grounding-based methods only reward final-answer correctness, even when localization is unreliable.
SpatialCORE is a post-training framework that turns the model's own grounding confidence into a learning signal:
- A self-regulating spatial reward weights each predicted bounding box's match quality by coordinate-token confidence, encouraging reasoning from confidently localized objects.
- An answer gate ties grounding optimization to final-answer correctness.
It achieves SOTA among open-source and specialized spatial reasoning models across benchmarks and transfers zero-shot to unseen distributions. Code is open-sourced.
More from Research
- Hugging Face open-sources Tau, a readable terminal coding agent built to teach — mervenoyann · 2026-10-01
- Researcher: new models solve old problems, but learning theory lacks predictive conjectures — brianryhuang · 2026-10-01
- ParallelPilot paper: 63% higher throughput for parallel coding agents — erichorvitz · 2026-10-01
- Multi-harness RL guide: LFM2.5 jumps 42% to 54% with 31% fewer tool calls — SergioPaniego · 2026-10-01
- LATENT wins IROS 2026 award: humanoid robots rally at human level from imperfect motion data — chris_j_paxton · 2026-10-01
- Bocconi paper: teach causal reasoning in the age of LLMs — daveholtz · 2026-10-01