The Seriality Gap in Video Models

mike64_t · x · 2026-07-17

This repost references a discussion on whether video models can act as world simulators, proposing a highly specific test: asking the model to predict the outcome of multiple balls colliding sequentially.

The conclusion is that standard video diffusion models break down significantly when the dependency chain grows too long, and simply throwing more compute at the problem won't fix it. The thread defines this phenomenon as the seriality gap—the model's fundamental deficiency in continuous, multi-step, serial causal reasoning.

Related event: Video Diffusion Models Expose Seriality Gap(2 posts)→

Original post →

More from Multimodal

Multimodal channel →