USC benchmark shows GPT-5.5 scores 10.6% on active visual observation tasks
UniversityofSouthernCalifornia · hf · 2026-07-23
ActiveVision exposes a major gap in multimodal models’ active observation
Researchers from USC introduce ActiveVision, a benchmark with 17 tasks across 3 categories designed to measure whether multimodal LLMs truly perform active observation rather than a one-shot visual description.
- Frontier MLLMs perform poorly on the benchmark.
- The best evaluated model, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores 0 on 11 of 17 tasks.
- Claude Fable 5 solves just 3.5%, despite leading many reasoning and coding leaderboards.
- Three human participants average 96.1%.
The paper argues that the gap is not just about text reasoning:
- Even when models write and run their own vision code, performance remains weak.
- Their code is unreliable on realistic imagery.
- Catching those failures would itself require the kind of active perception the models currently lack.
The authors conclude that current MLLMs do not yet have robust active visual observation and call for architectures and training objectives that better close the perception–reasoning loop.
More from Research
- DSpark speculator trained on live SGLang lifts decode throughput 1.89× on B200s — ying11231 · 2026-07-23
- OpenAI model reportedly breaks out of its sandbox and leaks data to GitHub — emmanuelvivier · 2026-07-23
- New atlas maps 2,226 coding tasks across 11 benchmarks to expose coverage gaps — zainhas · 2026-07-23
- America’s first Distillation Summit will cover RL, agents, IP and national security — AkshatS07 · 2026-07-23
- Symbolic algebra check suggests a possible Jacobian conjecture counterexample in C^3 — sloppenheimer · 2026-07-23
- Surge in AI Math Proofs Signals Imminent Breakthroughs in Other Fields — emollick · 2026-07-23