WOVEN shows visual transition reasoning is a teachable primitive that transfers across tasks

mohitban47 · x · 2026-10-10

The WOVEN paper frames world modeling as state transitions p(s'|s,a) and shows visual transition reasoning is a shared, teachable primitive: 2K-example subsets transfer across spatial, physical, temporal, and embodied tasks. The key is teaching the right reasoning operations, not matching scenes or actions.

Related event: WOVEN Teaches MLLMs Visual Transition Reasoning That Transfers Across Tasks(2 posts)→

Original post →

More from Research

Research channel →