NVIDIA's Spatial-IQ Benchmark Exposes Multimodal Models' Flaws in 3D Reasoning

NVIDIAAI · x · 2026-08-01

NVIDIA Research introduced Spatial-IQ, a diagnostic benchmark designed to evaluate 3D spatial reasoning capabilities. Tests show that humans achieve 82.1% accuracy in object counting tasks (including hidden objects), while the best off-the-shelf multimodal models score only 17.7%. By breaking spatial reasoning into 9 sub-tasks and applying targeted training, researchers successfully improved Qwen2.5-VL-32B's counting accuracy from 2.9% to 62.6%.

Related event: NVIDIA and Yale Introduce Spatial-IQ Benchmark(2 posts)→

Original post →

More from Models

Models channel →