Study: Video Diffusion Models Lack Scalable Sequential Compute

marksaroufim · x · 2026-07-20

A study titled "The Seriality Gap in Video Diffusion Models" points out that although video models are often called "world simulators," standard video diffusion models fail when handling long chains of dependent events (such as consecutive multi-ball collisions).

The research found that the model's denoising steps are essentially a form of recurrent computation rather than sequential computation, meaning they cannot improve reasoning capabilities simply by increasing compute. This implies that merely throwing more compute at the problem will not solve the deficits of video models in complex physical reasoning.

Original post →

More from Multimodal

Multimodal channel →