A new pixel-space benchmark says image generators can reason spatially, but text VLMs still lead

OmniAI-ZJU · hf · 2026-07-24

What the paper proposes

ProVisE (Protocolized Visual Evaluation) is a benchmark-agnostic framework for evaluating spatial cognition in image-generation models using visual answers instead of forcing text or coordinates.

Why it exists

Many spatial benchmarks ask models to output coordinates, multiple-choice options, or text, which creates an interface mismatch for image-generation models that can express answers directly in pixel space.

Main components

Findings

Takeaway

The work argues that evaluation format matters: pixel-space expression can expose strengths that text-only tests hide.

Original post →

More from Multimodal

Multimodal channel →