ByteDance's GST-Bench: VLMs score 42.68 vs human 79.08 on global spatial awareness
ByteDance-Seed · hf · 2026-08-07
ByteDance Seed introduces GST-Bench, a VQA benchmark for global spatial intelligence in video, with human-verified questions from 6,790 min of synthetic video. Evaluating 22 SOTA VLMs, the best zero-shot model scores 42.68, far below human 79.08. Models show strong local understanding but fail to consolidate long-horizon observations into a globally consistent scene representation. GST-Train dataset is also provided.
More from Research
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24
- Study: Agents read instructions/notes 60.5% of the time, rarely touch API docs — dair_ai · 2026-08-24