Frontier VLMs Collapse on Jigsaw Puzzles: New Benchmark Reveals Geometric Reasoning Cliff
kwangmoo_yi · x · 2026-08-01
A new paper introduces JigShape, a benchmark designed to evaluate the visual-geometric reasoning of Vision-Language Models (VLMs) using tab-and-blank interlocking jigsaw pieces to eliminate ground truth ambiguity.
Key findings from testing 95K instances reveal significant model limitations:
- Poor Zero-Shot Performance: Zero-shot VLMs largely lack geometric reasoning. Only GPT-5.5 slightly exceeds the random baseline on 4x4 puzzles, while all other frontier models perform at chance level.
- Scaling Cliff: While supervised fine-tuning achieves over 97% accuracy on 4x4 grids, all models collapse as grid density increases. Performance drops to near-random on 8x8 grids and falls below 5% on 12x12 puzzles.
This suggests that current architectures struggle to maintain consistent constraint satisfaction as the number of pieces scales up.
Related event: JigShape Benchmark Exposes Geometric Reasoning Flaws in VLMs(2 posts)→
More from Research
- Mouse Embryo Cell Lineage Reconstructed with DNA Typewriter, Over 1M Nodes — anshulkundaje · 2026-08-01
- Paper: Predictable Causal Architecture Exists Prior to the Origin of Life — drmichaellevin · 2026-08-01
- RLHF Data as the Core of the AI Economic Loop: Quality Data Design Determines Model Survival — herbiebradley · 2026-08-01
- Stanford Scholars Explore Generative and Agentic AI for Drug Discovery in Circulation — james_y_zou · 2026-08-01
- APEX-Accounting Benchmark: 58% Tasks Unsolved, Claude Fable 5 Takes the Lead — EdwardSun0909 · 2026-08-01
- Detecting Text Presence in Images: Best Architectures for Binary Classification — Relative-Pace-2923 · 2026-08-01