Inferring Physical Properties via Motion Probes
新智元 · wechat · 2026-07-16
PhyMAGIC aims to solve the issue of "a single image being insufficient to understand physical properties." Instead of directly guessing object parameters, it generates targeted motion probe videos, then uses a Vision-Language Model (VLM) to extract physical evidence from these motions and iteratively refine its judgments.
Methodology Highlights
- Uses an image-to-video model to generate motion probes like free fall, collision, squeezing, rotation, and side-pushing.
- Extracts key frames from the video for the VLM to infer parameters such as material, mass, density, Young's modulus, Poisson's ratio, and yield stress, estimating confidence levels for each.
- Automatically rewrites motion commands for low-confidence parameters to generate more targeted evidence, forming an "observe-infer-re-observe" closed loop.
- Combines physical parameters and simulator execution parameters into HPP, integrating 3D Gaussian Splatting with MPM simulations to render dynamic 3D objects viewable from multiple angles.
Results and Limitations
- Tested on PhysGaussian, PhysGen, and internet single-image scenarios, PhyMAGIC outperformed multiple baseline methods in CLIP similarity and Image-Motion-FID.
- 61 participants also favored it in ratings for physical plausibility and text consistency.
- However, it remains limited by single-image 3D reconstruction quality, VLM physical parameter estimation accuracy, and insufficient support for complex joint structures and multi-object interactions.
More from Multimodal
- Storyboard-first workflows are making AI dance videos and influencers more consistent — aftahi_ai · 2026-07-22
- Interactive video should be judged by responsiveness, not just frame quality — Soggy_Limit8864 · 2026-07-22
- Runpod MCP and Claude help spin up image and video generation workflows — 802high · 2026-07-22
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22
- A physics reward can improve video generation without creating a real physics engine — Dapper-Drawer4546 · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22