Video-model activations linearly expose position and predict collisions early
mathemagic1an · x · 2026-07-29
The thread argues that video models can encode physical state far more explicitly than they appear to from the outside.
- X/Y coordinates are linearly decodable from activations.
- Deeper layers can anticipate collisions several steps ahead.
- That suggests the model is not just interpolating pixels, but forming higher-level representations of objects and their relationships.
- The author uses this as motivation for the Jacobian-lens experiment that turns those internal directions into controllable motion.
More from Multimodal
- AI video turns a kidnapping premise into a self-duplication joke — umesh_ai · 2026-07-29
- Claude gaming art leaps: year-over-year comparison is stunning — iamfakhrealam · 2026-07-29
- SDXL Image Gen: How to Build Complex Multi-Character POV Interactions — ZeHirMan · 2026-07-29
- Hailuo AI’s new video model reportedly handles up to 12 image, video, and audio references — aziz4ai · 2026-07-29
- PDD accelerates image and video diffusion by predicting multiple denoising steps at once — ArashVahdat · 2026-07-29
- A weekly AI roundup packs robot MMA, FLUX 3 and an app-vs-art debate — PurzBeats · 2026-07-29