Dissecting Video Models with Jacobian Lens: 4M Params Emerge Causal World Model
mathemagic1an · x · 2026-07-29
The author explores whether modern video models truly understand physics or are merely stochastic pixel parrots.
Using a simple physics simulation (collisions, wall bounces, scoring), the author trained a 4M parameter pixel transformer to predict the next frame. Surprisingly, with just $10 of compute, the model demonstrated exceptional predictive capabilities.
- Emergent Physical Representation: The model isn't just memorizing the training set; it forms a higher-level representation capturing entities and their relationships, correctly deducing physical outcomes in unseen scenarios.
- Applying the Jacobian Lens: Using Anthropic's J-lens, researchers could linearly decode precise X/Y coordinates directly from the model's activations. Deeper layers even showed the model "anticipating" collisions several steps in advance.
- Playable Causal World Model: The model naturally contains causal directions for motion. By writing to specific directions in the activation space and mapping them to keyboard keys, the video model can literally be played like an interactive video game.
More from Fun
- Workplace meme turns HR into an “AI fantasy inside” room — Symbiot10000 · 2026-07-29
- “Home Model: Slop Fiction” leans into AI-generated absurdity — serialchilla91 · 2026-07-29
- AI video turns a kidnapping premise into a self-duplication joke — umesh_ai · 2026-07-29
- Eyecandy Robotics pitches robots as characters, not just machines — paulfinneyx · 2026-07-29
- A quantum-sensor poster turns magnetic navigation into techno-romantic art — paulfinneyx · 2026-07-29
- Claude gaming art leaps: year-over-year comparison is stunning — iamfakhrealam · 2026-07-29