Video models reveal a “Physics Emergence Zone” where motion direction becomes readable
mathemagic1an · x · 2026-07-29
A new interpretability study examines how large video transformers represent physics internally, using probing, subspace geometry, patch-level decoding, and attention ablations.
- The paper identifies a sharp intermediate-depth “Physics Emergence Zone” where physical variables become accessible.
- Speed and acceleration appear from early layers, while motion direction only becomes accessible in the emergence zone.
- Motion direction is encoded in a high-dimensional circular geometry, not as a clean factored variable.
- The authors conclude that video models do not store physics like a classical engine; instead they use distributed representations that are still sufficient for accurate prediction.
More from Research
- An AI digest scans 92 journals every week and turns them into one RSS feed — Afinetheorem · 2026-07-29
- A weekly PDB-synced leaderboard tracks open cofolding models — rishabh16_ · 2026-07-29
- Agentic AI Summit sets robotics and world models session with Sergey Levine and Jim Fan — dawnsongtweets · 2026-07-29
- New guidelines say GenAI still needs human oversight in systematic literature reviews — DrDatta_AIIMS · 2026-07-29
- Wharton labs open-sources AIBO for statistically valid AI behavior experiments — emollick · 2026-07-29
- A Unity RPG prompt pushes multi-agent coding loops to absurd AAA perfection — chongdashu · 2026-07-29