GUI-Primitives Benchmark Reveals VLM Spatial Reasoning Failures

usc-isi · hf · 2026-08-28

USC ISI released GUI-Primitives, a benchmark designed to diagnose spatial reasoning failures in vision-language models within GUI grounding tasks. The study reveals that models primarily fail to localize interface elements rather than misunderstanding spatial relations. Marking candidate elements before selection substantially improves performance.

Original post →

More from Research

Research channel →