Video-model activations linearly expose position and predict collisions early

mathemagic1an · x · 2026-07-29

The thread argues that video models can encode physical state far more explicitly than they appear to from the outside.

Related event: Research Shows 4M-Parameter Video Models Emerge Physical World Representations(4 posts)→

Original post →

More from Multimodal

Multimodal channel →