A new pixel-space benchmark says image generators can reason spatially, but text VLMs still lead
OmniAI-ZJU · hf · 2026-07-24
What the paper proposes
ProVisE (Protocolized Visual Evaluation) is a benchmark-agnostic framework for evaluating spatial cognition in image-generation models using visual answers instead of forcing text or coordinates.
Why it exists
Many spatial benchmarks ask models to output coordinates, multiple-choice options, or text, which creates an interface mismatch for image-generation models that can express answers directly in pixel space.
Main components
- A protocol-constrained visual-answer format that can be parsed into structured predictions
- An Agentic builder that creates and validates protocols for new benchmarks
- SpatialGen-Bench, a diagnostic benchmark with 470 samples across 14 spatial subtasks, four capability levels, and multiple answer forms
Findings
- Image-generation models can be competitive when spatial answers are externalized directly in pixel space.
- Text-output VLMs still have a clear edge in compositional spatial reasoning.
- The framework was validated on six external spatial benchmarks.
Takeaway
The work argues that evaluation format matters: pixel-space expression can expose strengths that text-only tests hide.
More from Multimodal
- Game-native world model adds 108M-frame dataset and explicit state supervision — Scobleizer · 2026-07-24
- A creator built a 4-minute animated short with FLUX.1 Dev and LTX-Video 2.3 — developervkmp · 2026-07-24
- VCSD boosts Qwen3-VL on ViRL39K without external teachers or evidence — UMCP · 2026-07-24
- Higgsfield’s MCP turns Claude into a no-code video studio with one URL — socialwithaayan · 2026-07-24
- A Claude plus video-model workflow can now turn one sentence into a 10-minute documentary — PrajwalTomar_ · 2026-07-24
- New Midjourney style shared with exact prompt settings and example images — azed_ai · 2026-07-24