Zhejiang Univ Introduces ProVisE: Evaluating Spatial Cognition via Pixels

机器之心 · wechat · 2026-08-08

The OmniAI team at Zhejiang University introduced the ProVisE framework, shifting away from traditional coordinate or text-based evaluations of AI spatial cognition by having generative image models directly "draw" their answers.

ProVisE Framework

SpatialGen-Bench

The team built a comprehensive benchmark covering 14 sub-tasks across four levels: perception, understanding, reasoning, and interaction.

Key Findings

Original post →

More from Multimodal

Multimodal channel →