Inferring Physical Properties via Motion Probes

新智元 · wechat · 2026-07-16

PhyMAGIC aims to solve the issue of "a single image being insufficient to understand physical properties." Instead of directly guessing object parameters, it generates targeted motion probe videos, then uses a Vision-Language Model (VLM) to extract physical evidence from these motions and iteratively refine its judgments.

Methodology Highlights

Results and Limitations

Original post →

More from Multimodal

Multimodal channel →