Zhejiang Univ Introduces ProVisE: Evaluating Spatial Cognition via Pixels
机器之心 · wechat · 2026-08-08
The OmniAI team at Zhejiang University introduced the ProVisE framework, shifting away from traditional coordinate or text-based evaluations of AI spatial cognition by having generative image models directly "draw" their answers.
ProVisE Framework
- Visual Protocols: Models are guided to mark targets, generate depth maps, or draw trajectories directly on images. A parser then converts these visual answers into structured data for scoring.
- Agentic Builder: An automated framework that scans task inputs and scoring rules to construct and validate corresponding "generation-parsing" protocols.
SpatialGen-Bench
The team built a comprehensive benchmark covering 14 sub-tasks across four levels: perception, understanding, reasoning, and interaction.
Key Findings
- Generative Strengths: Generative models exhibit stronger "spatial intuition" in tasks requiring direct visual perception, showing significant complementarity with text-based models.
- VLM Strengths: Text-output VLMs maintain an edge in tasks requiring abstract reasoning, such as counting or geometric feasibility checks.
More from Multimodal
- Hands-on with Grok Imagine 2.0: Precise Infographic Edits with Auto Color-Coding — AI_Andrew · 2026-08-08
- Testing Flux 3: Generating 'Found Footage' Analog Horror Videos — VIV-AF-D · 2026-08-08
- Testing Minimax H3 Character Swap: Impressive Expression Tracking — MIHAWKJR007 · 2026-08-08
- 324 Image Prompts and 42 AI Video Camera Movements Released — emmanuelvivier · 2026-08-08
- Testing H3 Text-to-Video on 'Mr. Robot': Stunning Atmosphere But Inconsistent Likeness — Positive_Writing_883 · 2026-08-08
- ComfyUI video gen issue: Batman's fins morph into feather-like flaps—how to keep consistency? — Guyserbun007 · 2026-08-08