USC benchmark shows GPT-5.5 scores 10.6% on active visual observation tasks
UniversityofSouthernCalifornia · hf · 2026-07-23
ActiveVision exposes a major gap in multimodal models’ active observation
Researchers from USC introduce ActiveVision, a benchmark with 17 tasks across 3 categories designed to measure whether multimodal LLMs truly perform active observation rather than a one-shot visual description.
- Frontier MLLMs perform poorly on the benchmark.
- The best evaluated model, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores 0 on 11 of 17 tasks.
- Claude Fable 5 solves just 3.5%, despite leading many reasoning and coding leaderboards.
- Three human participants average 96.1%.
The paper argues that the gap is not just about text reasoning:
- Even when models write and run their own vision code, performance remains weak.
- Their code is unreliable on realistic imagery.
- Catching those failures would itself require the kind of active perception the models currently lack.
The authors conclude that current MLLMs do not yet have robust active visual observation and call for architectures and training objectives that better close the perception–reasoning loop.
Related event: ActiveVision Benchmark Reveals Frontiers in Vision Models Lag Behind Humans(3 posts)→
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11