NVIDIA提出SparDA架构:解码提速1.7倍,长文本推理更准

HankYeomans · x · 2026-08-13

NVIDIA 研究团队提出了名为 SparDA(Sparse Decoupled Attention)的新型 Transformer 架构,旨在解决长上下文大模型推理的瓶颈。

核心改进与效果:

工程细节:

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →