Dissecting Video Models with Jacobian Lens: 4M Params Emerge Causal World Model

mathemagic1an · x · 2026-07-29

The author explores whether modern video models truly understand physics or are merely stochastic pixel parrots.

Using a simple physics simulation (collisions, wall bounces, scoring), the author trained a 4M parameter pixel transformer to predict the next frame. Surprisingly, with just $10 of compute, the model demonstrated exceptional predictive capabilities.

Related event: Research Shows 4M-Parameter Video Models Emerge Physical World Representations(4 posts)→

Original post →

More from Fun

Fun channel →