ActiveVision Benchmark Reveals Frontiers in Vision Models Lag Behind Humans
A new benchmark called ActiveVision from USC tests multimodal models' ability to actively observe and seek visual evidence. The results reveal significant weaknesses in front-tier vision models, with GPT-5.5 scoring only 10.6% compared to a 96.1% human average, highlighting a major gap in active visual perception.
2026-07-23 ~ 2026-07-24 · 3 related posts
- USC benchmark shows GPT-5.5 scores 10.6% on active visual observation tasks — UniversityofSouthernCalifornia · 2026-07-23
- GPT-5.5 scores 10.6% on ActiveVision as humans hit 96.1% — Justgototheeffinmoon · 2026-07-24
- ActiveVision benchmark finds frontier vision models far behind humans at repeated visual reasoning — OfirPress · 2026-07-24