New benchmark reveals poor visual perception in AI models; none reach 60% accuracy
The Decoder · rss · 2026-08-15
Moonshot AI released PerceptionBench, a benchmark designed to isolate visual perception from logical reasoning in multimodal AI models. Results show that all frontier models perform poorly, with none exceeding 60% accuracy; GPT-5.6 Sol leads by a narrow margin. The study suggests that many supposed reasoning failures actually occur during the image reading stage.
More from Models
- Post Mocks Meta's Massive Spending as China's Qwen3 Shows Strong Performance — ccerrato147 · 2026-08-15
- Distilled gstack version RL-finetuned in Kimi K3 weights — vikvang1 · 2026-08-15
- DeepSeek Falls Behind? User Claims It Trails Anthropic, OpenAI, and Others — scaling01 · 2026-08-15
- Polymarket is betting on DeepSeek's next Pro model: 60% odds by Nov 30 — Polymarket · 2026-08-15
- DeepSeek Developing Flash Variants to Match 3T Model Coding Performance — bindureddy · 2026-08-15
- Qwen3.8-27B performance sparks buzz, netizens joke about Google's reaction — max_paperclips · 2026-08-15