JigShape Benchmark Exposes Geometric Reasoning Flaws in VLMs

A new research benchmark called JigShape reveals severe flaws in the visual geometric reasoning of frontier Vision Language Models (VLMs), with performance on 8x8 puzzle tasks being almost equivalent to random guessing.

2026-08-01 ~ 2026-08-01 · 2 related posts