Survey: memory mechanisms for autoregressive video generation
Harold Haodong Chen · hf · 2026-09-24
A comprehensive review frames memory as the fundamental bottleneck of autoregressive video generation: under bounded context and compute, entity identities and causal changes leave active context long before they stop mattering. The paper defines memory operationally and organizes the literature across five lenses — Forms, Functions, Operations, Learning, and Evaluation — then synthesizes open challenges including composable memory architectures, trustworthy state updating, self-rollout learning, and standardized evaluation benchmarks.
More from Multimodal
- Creator open-sources brushstroke animation workflow built on Claude Opus 5.5 — alejandroll10 · 2026-09-24
- Meta's 3-year AI arc: from Threads to the Muse video model — minchoi · 2026-09-24
- Meta Connect showcases team's realtime interaction work — EdwardSun0909 · 2026-09-24
- 14-minute fully AI-generated series "Bride Of The Atom" Episode 1 released — geekycheekypixels · 2026-09-24
- DS 4.1 Flash drives Blender 3D animation, hinting multimodal training is essential for visual art — bookwormengr · 2026-09-24
- Muse adds voice generation; Scale AI CEO quips 'time to Pixar larp' — alexandr_wang · 2026-09-24