GPT-5.5 scores 10.6% on ActiveVision as humans hit 96.1%

Justgototheeffinmoon · reddit · 2026-07-24

A new arXiv paper says frontier vision models still struggle on ActiveVision, a benchmark designed to require repeated visual perception rather than one-shot description.

The authors say the point is not just that a vision model failed a new benchmark, but that the failure is highly structured and the models cannot repair it by writing their own code.

Related event: ActiveVision Benchmark Reveals Frontiers in Vision Models Lag Behind Humans(3 posts)→

Original post →

More from Models

Models channel →