SpatialCORE turns grounding confidence into a learning signal for spatial reasoning, hitting SOTA

Rafi Ibn Sultan · hf · 2026-10-01

Spatial reasoning remains a weakness of large vision-language models; existing grounding-based methods only reward final-answer correctness, even when localization is unreliable.

SpatialCORE is a post-training framework that turns the model's own grounding confidence into a learning signal:

It achieves SOTA among open-source and specialized spatial reasoning models across benchmarks and transfers zero-shot to unseen distributions. Code is open-sourced.

Original post →

More from Research

Research channel →