USC benchmark shows GPT-5.5 scores 10.6% on active visual observation tasks

UniversityofSouthernCalifornia · hf · 2026-07-23

ActiveVision exposes a major gap in multimodal models’ active observation

Researchers from USC introduce ActiveVision, a benchmark with 17 tasks across 3 categories designed to measure whether multimodal LLMs truly perform active observation rather than a one-shot visual description.

The paper argues that the gap is not just about text reasoning:

The authors conclude that current MLLMs do not yet have robust active visual observation and call for architectures and training objectives that better close the perception–reasoning loop.

Original post →

More from Research

Research channel →