GPT-5.5 scores 10.6% on ActiveVision as humans hit 96.1%
Justgototheeffinmoon · reddit · 2026-07-24
A new arXiv paper says frontier vision models still struggle on ActiveVision, a benchmark designed to require repeated visual perception rather than one-shot description.
- GPT-5.5 scores 10.6% at its highest exposed reasoning-effort tier and gets 0 on 11 of 17 tasks.
- Claude Fable 5 reaches only 3.5%, despite topping many reasoning and coding leaderboards.
- Three humans average 96.1%.
The authors say the point is not just that a vision model failed a new benchmark, but that the failure is highly structured and the models cannot repair it by writing their own code.
Related event: ActiveVision Benchmark Reveals Frontiers in Vision Models Lag Behind Humans(3 posts)→
More from Models
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27
- “Opus 5” post lands as a rebenchmarking-at-scale AI joke — kalomaze · 2026-07-27
- Top models now write worse than a year ago, critic says — dbreunig · 2026-07-27
- MPT-30B radar charts became an unexpectedly controversial design choice — code_star · 2026-07-27
- Local Gemma 4 31B starts acting sarcastic and users cannot reproduce it — n0head_r · 2026-07-27
- Google’s Gemini 3.6 Flash could win by matching Sonnet quality at a lower cost — haider1 · 2026-07-27