NVIDIA's Spatial-IQ Benchmark Exposes MLLM 3D Counting Flaws (17.7% vs 82.1% Human Accuracy)

NVIDIAAI · x · 2026-08-01

NVIDIA Research, in collaboration with Yale University, introduced Spatial-IQ, a diagnostic benchmark designed to evaluate the 3D spatial cognition of Multimodal Large Language Models (MLLMs).

The benchmark decomposes 3D object counting into nine hierarchical perceptual and cognitive sub-tasks, such as counting columns, layers, and inferring hidden supporting blocks. Experiments reveal that humans achieve 82.1% accuracy on these tasks, while the best off-the-shelf multimodal model scores only 17.7%.

The research highlights that current models lack the human-like ability to hierarchically decompose spatial structures. However, fine-tuning models on the Spatial-IQ dataset with chain-of-thought (CoT) supervision based on this proposed hierarchy can significantly improve their spatial reasoning capabilities.

Related event: NVIDIA and Yale Introduce Spatial-IQ Benchmark(2 posts)→

Original post →

More from Models

Models channel →