ActiveVision Benchmark Reveals Frontiers in Vision Models Lag Behind Humans

A new benchmark called ActiveVision from USC tests multimodal models' ability to actively observe and seek visual evidence. The results reveal significant weaknesses in front-tier vision models, with GPT-5.5 scoring only 10.6% compared to a 96.1% human average, highlighting a major gap in active visual perception.

2026-07-23 ~ 2026-07-24 · 3 related posts