SpaceCast-Bench: best VLM hits 58.0% on predictive spatial reasoning vs 87.2% human

zju · hf · 2026-10-09

ZJU's SpaceCast-Bench is the first benchmark directly evaluating predictive spatial reasoning—building scenes from observations and inferring unobserved outcomes—with 3,862 questions from 182 real scenes across 16 task types. The best of 21 models scores only 58.0% vs 87.2% human; spatial-specialized models are near chance. Bridge views prove critical, explicit 3D evidence beats generated outcome imagery, and fine-tuning lifts Qwen3-VL-4B from 34.0% to 65.7%.

Original post →

More from Models

Models channel →