Survey: memory mechanisms for autoregressive video generation

Harold Haodong Chen · hf · 2026-09-24

A comprehensive review frames memory as the fundamental bottleneck of autoregressive video generation: under bounded context and compute, entity identities and causal changes leave active context long before they stop mattering. The paper defines memory operationally and organizes the literature across five lenses — Forms, Functions, Operations, Learning, and Evaluation — then synthesizes open challenges including composable memory architectures, trustworthy state updating, self-rollout learning, and standardized evaluation benchmarks.

Original post →

More from Multimodal

Multimodal channel →