Alibaba DAMO unveils WorldAttention for efficient interactive video world models
Alibaba-DAMO-Academy · hf · 2026-09-30
Alibaba DAMO Academy 发布 WorldAttention,一套面向文本条件交互式视频世界模型的高效注意力架构,解决长时程生成中的历史上下文与显存瓶颈问题。核心设计:
- Hybrid Sparse Attention (HSA):线性全局注意力 + 头自适应稀疏注意力,替代传统滑窗机制造成的上下文丢失。
- Hierarchical KV Cache (HKV):将历史 KV 对按语义索引组织到多级存储页中,细粒度检索并控制 GPU 驻留,避免 KV cache 线性增长导致显存饱和。
- 配套定制 kernel 将理论效率转化为实际性能。
在 VBench-Long 和 InterVBench 上全面超越此前 SOTA,主体一致性分别达 0.9472 和 0.9668。该架构瞄准具身智能与基于仿真的规划所需的低延迟、长时长生成场景。
More from Multimodal
- AutoRef open-sourced: harness optimization for agentic multi-reference image generation — NunyaBuzor · 2026-09-30
- Claude Opus 5.5 video imagines the moment Girl with a Pearl Earring was painted — justin_hart · 2026-09-30
- Reddit user shares AI-generated sci-fi short film Erebus-9: The Hiveborn Archives — himeonisama · 2026-09-30
- Teaching Claude Opus to Annotate Speech Timing for ElevenLabs Voiceovers Proves Tricky — RileyRalmuto · 2026-09-30
- NUS's MaLiang-Harness exposes the Program-to-Visual gap in code-driven image/video generation — NationalUniversityofSingapore · 2026-09-30
- Thinking Reward Model: rubric-first scoring sets open-source visual generation reward SOTA — Xuehai Bai · 2026-09-30