PhiZero: A Physical World Model Driven by Physical Language
Shuyao Shang · hf · 2026-07-31
PhiZero introduces a novel physical world model built around "physical language"—a compact, discrete representation of world-state transitions. Unlike existing models that predict future videos directly in pixel space, leaving physical dynamics implicit, PhiZero learns to abstract predictive structures into physical language via self-supervision from in-the-wild videos.
Core Mechanism:
- Adopts a reason-then-render paradigm: it first infers future world evolution as a sequence of physical language, then renders these transitions into videos.
Key Results:
- Validates its ability to model physically coherent world evolution across generation and understanding benchmarks.
- Demonstrates strong potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.
More from Research
- Meta Proposes OneShot Retrieval Framework, Deployed in Instagram — _reachsumit · 2026-07-31
- Google Proposes HA-MoE Architecture to Boost Cross-Content Ranking Fairness — _reachsumit · 2026-07-31
- Research: Enhancing Generative Recommendation with LLM-Derived Language Tags — _reachsumit · 2026-07-31
- Meta's ROCS Paradigm Boosts Recommendation Retrieval QPS Up to 3x — _reachsumit · 2026-07-31
- With Mandatory Reviewing, Low-Quality Reviews in AI Conferences Are No Longer Justifiable — Kwangryeol · 2026-07-31
- HiLaR: Optimizing LLM Recommendation Reasoning via Hierarchical RL — _reachsumit · 2026-07-31