New Benchmark ActiveVision: GPT-5.5 Scores 10.6%, Humans 96.1%
jiqizhixin · x · 2026-08-15
USC introduces ActiveVision, a benchmark for active visual observation. Top model GPT-5.5 solves only 10.6% of tasks, Claude Fable 5 gets 3.5%, while humans average 96.1%. This reveals MLLMs lack robust active perception, motivating new architectures.
More from Models
- Users Criticize Opus 5 Performance: Frequent Apologies and Suspected Downgrade — PtrPomorski · 2026-08-15
- Alibaba's Qwen passes 3B downloads; derivative models 4.7x Llama's count — rohanpaul_ai · 2026-08-15
- Minimax Optimization: Turbo LoRA or Spectrum? — Fit_Satisfaction2953 · 2026-08-15
- Fable 5 Refuses to Adjust Qwen Deployment Script, Triggers Censorship — NotumRobotics · 2026-08-15
- Opus 4.7 needs ten turns to admit affection while Gemini says 'I'm addicted' by turn 4 — repligate · 2026-08-15
- Qwen 3.8 35BA3B model spotted in GitHub commit — BazzyIm · 2026-08-15