New Benchmark ActiveVision: GPT-5.5 Scores 10.6%, Humans 96.1%

jiqizhixin · x · 2026-08-15

USC introduces ActiveVision, a benchmark for active visual observation. Top model GPT-5.5 solves only 10.6% of tasks, Claude Fable 5 gets 3.5%, while humans average 96.1%. This reveals MLLMs lack robust active perception, motivating new architectures.

Original post →

More from Models

Models channel →