USTC's PhysVista benchmark exposes wide gap between VLM visual recognition and physical understanding

ustc · hf · 2026-10-02

Researchers from USTC released PhysVista, a benchmark evaluating physical intelligence in vision-language models via a closed "perception-reasoning-assessment" cognitive loop. It jointly tests physical state perception, dynamics reasoning, and plausibility assessment; distinguishes event-level and scale-level reasoning; and includes both real-world and AI-generated videos to test whether models can spot physical inconsistencies in generative content. Experiments across diverse VLMs reveal substantial limitations in physical reasoning, highlighting a persistent gap between visual recognition and genuine physical understanding.

Original post →

More from Multimodal

Multimodal channel →