Google DeepMind 论文 Declarative Attention:让模型自己声明读哪段 KV cache

omarsar0 · x · 2026-09-04

Google DeepMind 及合作者发布论文 Declarative Attention,直击长上下文推理的核心浪费:模型每生成一个 token 都要重读全部 KV cache,即使实际只关注其中一小部分——100 万 token 对话中问一个细节,全局注意力层每步都要重读全部缓存。

所属事件:Declarative Attention:让语言模型自主声明读哪段 KV cache(5 条相关)→

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →