How video world models learn intuitive physics: V-JEPA analysis
soniajoseph_ · x · 2026-08-18
Sonia Joseph's talk explores how video world models represent and reason about intuitive physics, focusing on the emergence of physical reasoning within V-JEPA-style architectures. It also covers mechanistic interpretability for video models and PhD advice in the current AI climate.
More from Multimodal
- Seedance 2.5 demo: Vehicle evolves into sci-fi machine every few seconds — umesh_ai · 2026-08-18
- Video Model Evaluation Guide: Quantifying Generation Quality — Majumdar_Ani · 2026-08-18
- AI can now control camera precisely with a single line of code — jnack · 2026-08-18
- Krea 2 training deep-dive: data is everything, and never train on AI images — AI Engineer · 2026-08-18
- Which AI video tools are worth using in 2026: hands-on picks from Sora to Veo — jimmybobjoeflow · 2026-08-18
- Stable Audio 3.0 gets a DAW plugin and iterative web studio, free for commercial use — Stability AI · 2026-08-18