WOVEN: visual transition reasoning as a shared primitive lifts 22 of 26 benchmarks by up to 27.3 points

mohitban47 · x · 2026-10-10

New work WOVEN studies what makes world-model training for MLLMs actually transfer across tasks. The authors frame world modeling as state transitions p(s′ | s, a) and focus on visual transition reasoning — inferring what is unobserved in (s, a, s′) — asking whether it can be learned as a shared primitive benefiting multiple downstream tasks.

Key takeaways: it is teachable and transfers widely — subsets of only 2K examples collectively improve 22 of 26 downstream benchmarks by up to 27.3 points. They also propose a data-centric recipe: select supervision by the reasoning operation it teaches. Tasks covered include spatial, physical, temporal, and embodied reasoning.

Related event: WOVEN Teaches MLLMs Visual Transition Reasoning That Transfers Across Tasks(2 posts)→

Original post →

More from Research

Research channel →