PhiZero World Model Uses 'Physical Language' to Transfer Skills Across Robot Morphologies

青稞AI · wechat · 2026-08-08

Current video world models often predict the future directly in pixel space, prioritizing visual realism over physical accuracy. Researchers at CASIA propose PhiZero, which extracts compact, discrete "Physical Language" from massive real-world videos via self-supervised learning, decoupling what the world looks like from how it changes.

PhiZero adopts a "Reason-then-Render" paradigm: it first predicts a sequence of physical state changes, which a diffusion decoder then renders into video. This representation of state transitions is highly reusable, enabling impressive cross-embodiment transfer. For instance, motion and interaction logic from human or simulation videos can be transferred and rendered onto a Unitree G1 humanoid robot or a dexterous hand without paired training data.

Across multiple video generation and physics understanding benchmarks, PhiZero achieves leading results in physical consistency and causal reasoning, offering a new paradigm for embodied AI planning and cross-robot skill transfer.

Original post →

More from Embodied

Embodied channel →