Frontier VLMs Fail at Jigsaw Puzzles: JigShape Benchmark Tests Geometric Reasoning
_vztu · x · 2026-08-04
While training models on jigsaw puzzles improves object detection and 3D understanding, current frontier Vision-Language Models (VLMs) struggle significantly with solving them.
The team introduces JigShape, a benchmark designed to evaluate visual-geometric reasoning in VLMs. Unlike prior benchmarks that cut images into interchangeable rectangles, JigShape uses real interlocking puzzle pieces (with tabs and blanks). This enforces unique ground truth, eliminating the ambiguity of interchangeable patches like blank sky. The project also releases an open-source training dataset of 186k rows.
More from Research
- DAPD Paper Tackles Information Asymmetry in LLM Policy Distillation — _akhaliq · 2026-08-05
- Pruning Attention Layers Slashes Action Expert FLOPs by 94% in Robotics — mathildepapillo · 2026-08-05
- Probing VLM Medical Image Recognition with Silico Platform — mathildepapillo · 2026-08-05
- Goodfire launches Silico interpretability tool for researchers at $1,000/mo — deedydas · 2026-08-05
- Interpretability Analysis Optimizes Robotics Model, Cutting Compute by 40% — mathildepapillo · 2026-08-05
- Investigating Information Flow Between VLM and Action Experts in VLAs — mathildepapillo · 2026-08-05