Frontier VLMs Fail at Jigsaw Puzzles: JigShape Benchmark Tests Geometric Reasoning

_vztu · x · 2026-08-04

While training models on jigsaw puzzles improves object detection and 3D understanding, current frontier Vision-Language Models (VLMs) struggle significantly with solving them.

The team introduces JigShape, a benchmark designed to evaluate visual-geometric reasoning in VLMs. Unlike prior benchmarks that cut images into interchangeable rectangles, JigShape uses real interlocking puzzle pieces (with tabs and blanks). This enforces unique ground truth, eliminating the ambiguity of interchangeable patches like blank sky. The project also releases an open-source training dataset of 186k rows.

Original post →

More from Research

Research channel →