JigShape Benchmark: VLMs Fail at Visual-Geometric Reasoning

Shawn Li · hf · 2026-08-12

Introduces JigShape, a new jigsaw benchmark with interlocking pieces to evaluate VLMs. It reveals that vision-language models fail at geometric reasoning and suffer a sharp performance drop as puzzle size increases.

Original post →

More from Research

Research channel →