New benchmark reveals poor visual perception in AI models; none reach 60% accuracy

The Decoder · rss · 2026-08-15

Moonshot AI released PerceptionBench, a benchmark designed to isolate visual perception from logical reasoning in multimodal AI models. Results show that all frontier models perform poorly, with none exceeding 60% accuracy; GPT-5.6 Sol leads by a narrow margin. The study suggests that many supposed reasoning failures actually occur during the image reading stage.

Original post →

More from Models

Models channel →