ActiveVision Benchmark: Top VLMs Lag Humans by 9x in Active Visual Reasoning
机器之心 · wechat · 2026-08-05
A USC team released ActiveVision, a benchmark testing multimodal models' 'active observation' skills (e.g., global scanning, sequential traversal, attribute transfer).
Key Findings:
- Pure CoT Models Fail: Without tools, the best model scored only 10.6% versus 96.1% for humans—a 9x gap. Increasing reasoning strength barely helped, as models fail to form a closed loop of evidence gathering and verification.
- Coding Agents Still Fall Short: Tool-augmented agents (Fable5, GPT-5.5) improved scores (up to 50.6%) by turning 'seeing' into 'calculating,' but failed heavily on tasks requiring continuous tracking. They also took 12-15 minutes and cost several dollars per question, compared to 34 seconds for humans.
The benchmark uses procedurally generated, photorealistic images to prevent shortcut-taking, exposing severe bottlenecks in continuous visual perception.
More from Models
- Qwen Integrates with Cline: $4.99 Promo Offers Discounted API Access — Alibaba_Qwen · 2026-08-05
- User Test: Chinese LLM Claims Distilled Origins Without System Prompts — _TheWolfOfWalmart_ · 2026-08-05
- Over-Optimized AI Models: A Shared Regression in Instruction Following — spleentastic · 2026-08-05
- MiniMax Issues Takedown Warnings for Decensor LoRAs, Threatens License Revocation — Xto · 2026-08-05
- Qwen 3.8 Matches Fable 5 in Video Quality at 40% of the Cost — ChrisGPT · 2026-08-05
- GPT-5.6 Luna High Rumored to Offer Unlimited Cheap Usage; Devs Eye Browser Automation — petergyang · 2026-08-05