Qwen's new Flash model and Meituan LongCat both use N-gram to work around HBM shortage

bookwormengr · x · 2026-09-21

The author argues model architectures are shaped by available hardware: Qwen-3.8-Flash-Next also uses N-gram, a technique Meituan's LongCat lab independently invented. Labs without enough HBM work around it with techniques like N-gram and WideEP to reduce HBM needs — current architectures are not necessarily optimal.

Original post →

More from Infra

Infra channel →