ByteDance's GST-Bench: VLMs score 42.68 vs human 79.08 on global spatial awareness

ByteDance-Seed · hf · 2026-08-07

ByteDance Seed introduces GST-Bench, a VQA benchmark for global spatial intelligence in video, with human-verified questions from 6,790 min of synthetic video. Evaluating 22 SOTA VLMs, the best zero-shot model scores 42.68, far below human 79.08. Models show strong local understanding but fail to consolidate long-horizon observations into a globally consistent scene representation. GST-Train dataset is also provided.

Original post →

More from Research

Research channel →