JigShape Benchmark: VLMs Fail at Visual-Geometric Reasoning
Shawn Li · hf · 2026-08-12
Introduces JigShape, a new jigsaw benchmark with interlocking pieces to evaluate VLMs. It reveals that vision-language models fail at geometric reasoning and suffer a sharp performance drop as puzzle size increases.
More from Research
- Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training — _reachsumit · 2026-08-12
- Sequential Modality Dropout: A 4-Line Fix for Robust Multi-Modal Recommenders — _reachsumit · 2026-08-12
- Netflix's GenRec: LLM-Backed Recommendation Ranker Beats Production with 40x Less Data — _reachsumit · 2026-08-12
- NTCF: Tree Collaborative Filtering Framework with Curvature-Aware Propagation Depth — _reachsumit · 2026-08-12
- Model Merging Framework Shortens LLM Recommender Reasoning Traces by 24% — _reachsumit · 2026-08-12
- MIJSR Framework: Jointly Mining Multi-Interests for Search and Recommendation — _reachsumit · 2026-08-12