Frontier VLMs Collapse on Jigsaw Puzzles: New Benchmark Reveals Geometric Reasoning Cliff

kwangmoo_yi · x · 2026-08-01

A new paper introduces JigShape, a benchmark designed to evaluate the visual-geometric reasoning of Vision-Language Models (VLMs) using tab-and-blank interlocking jigsaw pieces to eliminate ground truth ambiguity.

Key findings from testing 95K instances reveal significant model limitations:

This suggests that current architectures struggle to maintain consistent constraint satisfaction as the number of pieces scales up.

Related event: JigShape Benchmark Exposes Geometric Reasoning Flaws in VLMs(2 posts)→

Original post →

More from Research

Research channel →