How video world models learn intuitive physics: V-JEPA analysis

soniajoseph_ · x · 2026-08-18

Sonia Joseph's talk explores how video world models represent and reason about intuitive physics, focusing on the emergence of physical reasoning within V-JEPA-style architectures. It also covers mechanistic interpretability for video models and PhD advice in the current AI climate.

Original post →

More from Multimodal

Multimodal channel →