ActiveVision benchmark finds frontier vision models far behind humans at repeated visual reasoning

OfirPress · x · 2026-07-24

ActiveVision is a new benchmark for testing whether models can repeatedly observe, reason, and seek new visual evidence instead of making a single glance decision.

The result suggests current vision models are still far from robust multi-step visual reasoning in interactive settings.

Related event: ActiveVision Benchmark Reveals Frontiers in Vision Models Lag Behind Humans(3 posts)→

Original post →

More from Research

Research channel →