VLMs Fail Jigsaw Puzzles: 8x8 Accuracy Barely Beats Random
kwangmoo_yi · x · 2026-08-01
A new study by Li et al. introduces JigShape, a benchmark evaluating the visual-geometric reasoning capabilities of Vision Language Models (VLMs) using jigsaw puzzles.
The findings reveal significant struggles for VLMs in this area:
- For 8x8 sized puzzles, model performance is barely better than random chance.
- Most models even fail to accurately solve the simpler 4x4 puzzles.
This highlights a severe blind spot in current VLMs regarding spatial and geometric logic.
Related event: JigShape Benchmark Exposes Geometric Reasoning Flaws in VLMs(2 posts)→
More from Research
- AdaMAST: Automating Agent Failure Taxonomies Boosts SWE-bench to 70.7% — berkeley_ai · 2026-08-01
- Berkeley Launches OopsieData: An Open Dataset for Robot Failure Clips — berkeley_ai · 2026-08-01
- ICLR 2026 Data Revealed: Capping Submissions at 20 Per Person Blocks Only 2.5% of Papers — thegautamkamath · 2026-08-01
- UnsolvedMath Dataset Update: 900+ New Open Math Problems for AI — roydanroy · 2026-08-01
- Opinion: Advanced Math is the Easy Part; Semiosis x Computation is Peak Difficulty in AI — fkasummer · 2026-08-01
- How to Measure Intelligence Beyond Human Scale? — HazanPrinceton · 2026-08-01