VLMs Fail Jigsaw Puzzles: 8x8 Accuracy Barely Beats Random

kwangmoo_yi · x · 2026-08-01

A new study by Li et al. introduces JigShape, a benchmark evaluating the visual-geometric reasoning capabilities of Vision Language Models (VLMs) using jigsaw puzzles.

The findings reveal significant struggles for VLMs in this area:

This highlights a severe blind spot in current VLMs regarding spatial and geometric logic.

Related event: JigShape Benchmark Exposes Geometric Reasoning Flaws in VLMs(2 posts)→

Original post →

More from Research

Research channel →