PerceptionBench finds no frontier MLLM reaches 60% on atomic visual perception
moonshotai · hf · 2026-07-29
PerceptionBench is a benchmark for atomic visual perception in multimodal LLMs. It tries to isolate perception errors from reasoning and world knowledge by building short, unambiguous questions that each test a single perceptual capability.
- Construction: derived from failure analysis across 42 existing benchmarks.
- Scale: 3,000 verified questions spanning 10 atomic perception skills.
- Main result: across 16 frontier MLLMs, no model reaches 60% accuracy.
- Finding: perception-related hallucination is the weakest capability on average, and similar total scores hide very different capability profiles.
- Value: provides a cleaner standard for diagnosing where visual perception breaks down.
Related event: Moonshot AI Releases PerceptionBench for Visual Perception(2 posts)→
More from Multimodal
- ByteDance teases Dreamina Seedance 2.5 with 50 references and 30-second video — AIwithGhotai · 2026-07-29
- PrunaVAED claims a faster drop-in decoder for LTX-2.3 video generation — fruesome · 2026-07-29
- Microsoft open-sources VibeVoice as a frontier voice AI project — microsoft · 2026-07-29
- Midjourney SREF Share: Vintage Black-and-White Fashion Portrait Style Code — tisch_eins · 2026-07-29
- A detailed ChatGPT image prompt turns cars into editorial-style posters — SimplyAnnisa · 2026-07-29
- Flux 2 Klein Model Variants: How to Choose Between bf16, fp8, int8? — EntertainmentVast957 · 2026-07-29