Qwen 新模型与美团 LongCat 均用 N-gram 架构绕开 HBM 短缺

bookwormengr · x · 2026-09-21

The author argues model architectures are shaped by available hardware: Qwen-3.8-Flash-Next also uses N-gram, a technique Meituan's LongCat lab independently invented. Labs without enough HBM work around it with techniques like N-gram and WideEP to reduce HBM needs — current architectures are not necessarily optimal.

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →