Rowmax-H15:B200 上注意力前向最高提速 25.8%

illinois · hf · 2026-09-30

研究者在 10 个冻结 decoder-only 模型(0.5B–72B)上系统研究预训练 Transformer 推理时 softmax 可以近似到什么程度。

发现:

提出:

实测(B200):FP8 注意力前向 causal 8K 提速 12.4%,non-causal 8K 提速 25.8%;causal 16K 每次前向板级能耗降 8.4%;BF16 路径 2K 下五模型困惑度仅升 0.091–0.492%。

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →