VDiff-Bench: 1,756-question benchmark shows frontier models fail at spot-the-difference

yixin_wan_ · x · 2026-09-10

Researchers introduce VDiff-Bench, a challenging benchmark for fine-grained image difference identification: 1,756 four-way multiple-choice questions over image pairs covering 10 change categories (position, motion, color, noise/resolution, texture, OCR/text, illumination, etc.), with ground-truth-conditioned negatives to make the task harder.

Tests on 11 state-of-the-art open- and closed-source MLLMs show fine-grained visual comparison remains brittle:

Paper on arXiv (2609.06245), dataset on Hugging Face. A contrast with frontier-science claims: models can do research, yet fail at a kids' spot-the-difference game.

Original post →

More from Models

Models channel →