A new pixel-space benchmark says image generators can reason spatially, but text VLMs still lead
OmniAI-ZJU · hf · 2026-07-24
What the paper proposes
ProVisE (Protocolized Visual Evaluation) is a benchmark-agnostic framework for evaluating spatial cognition in image-generation models using visual answers instead of forcing text or coordinates.
Why it exists
Many spatial benchmarks ask models to output coordinates, multiple-choice options, or text, which creates an interface mismatch for image-generation models that can express answers directly in pixel space.
Main components
- A protocol-constrained visual-answer format that can be parsed into structured predictions
- An Agentic builder that creates and validates protocols for new benchmarks
- SpatialGen-Bench, a diagnostic benchmark with 470 samples across 14 spatial subtasks, four capability levels, and multiple answer forms
Findings
- Image-generation models can be competitive when spatial answers are externalized directly in pixel space.
- Text-output VLMs still have a clear edge in compositional spatial reasoning.
- The framework was validated on six external spatial benchmarks.
Takeaway
The work argues that evaluation format matters: pixel-space expression can expose strengths that text-only tests hide.
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11