Qwen's new Flash model and Meituan LongCat both use N-gram to work around HBM shortage
bookwormengr · x · 2026-09-21
The author argues model architectures are shaped by available hardware: Qwen-3.8-Flash-Next also uses N-gram, a technique Meituan's LongCat lab independently invented. Labs without enough HBM work around it with techniques like N-gram and WideEP to reduce HBM needs — current architectures are not necessarily optimal.
More from Infra
- SGLang's hicache: use an L3 storage cache to keep KV cache alive across local model swaps — TheZachMueller · 2026-09-22
- Egypt's AI Ecosystem Hits Production Scale With $400M Data Center, 10x NVIDIA Learner Growth — nordicinst · 2026-09-22
- RTX Pro 6000 vs a used 3090 vs cloud rental: the LoRA training math — big-in-jap · 2026-09-21
- ComfyUI GPU rental showdown: Modal's 35s cold starts and free 1TiB beat RunPod — ronalder100 · 2026-09-21
- Meta partners with Arm on Arm AGI CPU, its first AI-era data center CPU — bookwormengr · 2026-09-21
- Starlink is becoming core infrastructure for rural education across Latin America — XFreeze · 2026-09-21