SpaceCast-Bench: best VLM hits 58.0% on predictive spatial reasoning vs 87.2% human
zju · hf · 2026-10-09
ZJU's SpaceCast-Bench is the first benchmark directly evaluating predictive spatial reasoning—building scenes from observations and inferring unobserved outcomes—with 3,862 questions from 182 real scenes across 16 task types. The best of 21 models scores only 58.0% vs 87.2% human; spatial-specialized models are near chance. Bridge views prove critical, explicit 3D evidence beats generated outcome imagery, and fine-tuning lifts Qwen3-VL-4B from 34.0% to 65.7%.
More from Models
- Open-Source NSFW Classifier Blue-Eye Hits 88.9%, Beats AWS and Google Vision — Mundane_Toe_8074 · 2026-10-09
- Every model looked bad in my eval — the bug was my answer key, not the models — jgarg27 · 2026-10-09
- Ai2: the next Olmo model is already training, fully open-model commitment unchanged — sewon__min · 2026-10-09
- Whistle: an open 16.9 MB speech-to-text model that runs on CPU with 11 ms first token in 7 languages — solyarisoftware · 2026-10-09
- Nous Research's typo 'Hemres Agnet' looks like a Hermes Agent teaser — Teknium · 2026-10-09
- Study: Outdated Gemini 2.5 Advice Rated on Par With Doctors in Urgent Care, No Safety Issues — emollick · 2026-10-09