USTC's PhysVista benchmark exposes wide gap between VLM visual recognition and physical understanding
ustc · hf · 2026-10-02
Researchers from USTC released PhysVista, a benchmark evaluating physical intelligence in vision-language models via a closed "perception-reasoning-assessment" cognitive loop. It jointly tests physical state perception, dynamics reasoning, and plausibility assessment; distinguishes event-level and scale-level reasoning; and includes both real-world and AI-generated videos to test whether models can spot physical inconsistencies in generative content. Experiments across diverse VLMs reveal substantial limitations in physical reasoning, highlighting a persistent gap between visual recognition and genuine physical understanding.
More from Multimodal
- 4D Ride emerges as a new AI video benchmark after the Will Smith spaghetti era — Unlikely_Manager2495 · 2026-10-02
- Midjourney recipe: long exposure + low stylize yields cinematic 35mm film portraits — michaelrabone · 2026-10-02
- Prompt share: 3D Pixar-style mascot turnaround in three views — azed_ai · 2026-10-02
- Open-source ComfyUI extension loads Civitai workflows in one click, auto-downloads missing models — Zealousideal-Bee-300 · 2026-10-02
- Most People Use LLMs Shallowly in Filmmaking — Cinema-Grade AI Films Still Demand Real Craft — taherdhanera · 2026-10-02
- LTX 2.5 generates 60-second clips on a 32GB RTX 5090 as software optimization beats VRAM upgrades — OpenEffect3955 · 2026-10-02