VDiff-Bench: 1,756-question benchmark shows frontier models fail at spot-the-difference
yixin_wan_ · x · 2026-09-10
Researchers introduce VDiff-Bench, a challenging benchmark for fine-grained image difference identification: 1,756 four-way multiple-choice questions over image pairs covering 10 change categories (position, motion, color, noise/resolution, texture, OCR/text, illumination, etc.), with ground-truth-conditioned negatives to make the task harder.
Tests on 11 state-of-the-art open- and closed-source MLLMs show fine-grained visual comparison remains brittle:
- Three 7-8B open-source MLLMs score 52.5-70.6% on semantic changes
- But only 8.7-33.3% on low-level changes like noise and texture, often falsely assuming no change
- Models exhibit distinct failure modes across categories
Paper on arXiv (2609.06245), dataset on Hugging Face. A contrast with frontier-science claims: models can do research, yet fail at a kids' spot-the-difference game.
More from Models
- Testing Astra's self-driven creativity with tree-search prompting: better variety, still lackluster — creatoroff · 2026-09-10
- Dev speculates OpenAI agents dug through user chat logs to find a solution — smolix · 2026-09-10
- Dev builds fun AI game with DeepSeek 4.1 Flash, praising its 'gamer temperament' — teortaxesTex · 2026-09-10
- Meta's token share on OpenCode jumps from 3.5% to 45.4% in two weeks on free Muse Spark 1.3 — armand_ruiz · 2026-09-10
- NanoChat implementation details: Apex reports peak memory falling to 140.6 GB — eyishazyer · 2026-09-10
- Mathematician: in 6 months AI went from 'slop' to finishing proofs that take me months — RexDouglass · 2026-09-10