WOVEN: visual transition reasoning as a shared primitive lifts 22 of 26 benchmarks by up to 27.3 points
mohitban47 · x · 2026-10-10
New work WOVEN studies what makes world-model training for MLLMs actually transfer across tasks. The authors frame world modeling as state transitions p(s′ | s, a) and focus on visual transition reasoning — inferring what is unobserved in (s, a, s′) — asking whether it can be learned as a shared primitive benefiting multiple downstream tasks.
Key takeaways: it is teachable and transfers widely — subsets of only 2K examples collectively improve 22 of 26 downstream benchmarks by up to 27.3 points. They also propose a data-centric recipe: select supervision by the reasoning operation it teaches. Tasks covered include spatial, physical, temporal, and embodied reasoning.
Related event: WOVEN Teaches MLLMs Visual Transition Reasoning That Transfers Across Tasks(2 posts)→
More from Research
- Antoine Chaffin: backbones are strong enough that specialized synthetic data isn't required yet — antoine_chaffin · 2026-10-10
- Strong backbones plus light fine-tuning beat synthetic data, says researcher whose model tops benchmarks — antoine_chaffin · 2026-10-10
- Maxime Rivest asks: does turning a backbone into a product model require any training at all? — MaximeRivest · 2026-10-10
- Simulator + lab-in-the-loop is the standard for protein optimization, says Kidger — PatrickKidger · 2026-10-10
- Meta paper shows byte-level models beat tokenized ones given enough training compute — alex_verem · 2026-10-10
- Antoine Chaffin: you don't even need synthetic data, backbones are already very strong — antoine_chaffin · 2026-10-10