ICML oral Motive: first motion attribution framework for video generation, 74.1% win rate

cindy_x_wu · x · 2026-09-04

The ICML 2026 oral 'Motion Attribution for Video Generation' (Xindi Wu, Antonio Torralba, Sanja Fidler, Jonathan Lorraine, et al.) presents Motive, the first gradient-based framework attributing motion (rather than visual appearance) in video generation models. Using motion-weighted loss masks to isolate temporal dynamics, Motive identifies which fine-tuning clips improve or degrade motion and guides data curation, improving motion smoothness and dynamic degree on VBench with a 74.1% human preference win rate over the base model.

Related event: ICML oral paper on motion attribution for video generation released(2 posts)→

Original post →

More from Multimodal

Multimodal channel →