ActiveVision Benchmark: Top VLMs Lag Humans by 9x in Active Visual Reasoning

机器之心 · wechat · 2026-08-05

A USC team released ActiveVision, a benchmark testing multimodal models' 'active observation' skills (e.g., global scanning, sequential traversal, attribute transfer).

Key Findings:

The benchmark uses procedurally generated, photorealistic images to prevent shortcut-taking, exposing severe bottlenecks in continuous visual perception.

Original post →

More from Models

Models channel →