Alibaba's Qwen3.8 with 51B N-gram embeddings now available on SGLang
Alibaba_Qwen · x · 2026-08-26
Alibaba announced that Qwen3.8-Flash-Next is ready for deployment via SGLang. The architecture features a 125B main model paired with 51B N-gram embeddings, activating only 6B parameters per token.
Key highlights:
- N-gram embeddings: Scale capacity with minimal extra compute per token, residing in host memory with async prefetch to save GPU VRAM.
- GDN + QSA hybrid attention: Balances memory efficiency and precise retrieval for long-horizon tasks.
- Gated Residual: Expands information pathways between layers from 1 to 4 lanes.
The model was trained using the Muon optimizer.
More from Infra
- OpenAI plans world's largest data center in Ohio, requiring more power than all state homes — bennash · 2026-08-26
- Vercel AI Gateway adds asynchronous video generation with Webhook support — cramforce · 2026-08-26
- Knowledge Compressor cuts doc tokens in half to save Agent context costs — mariorod1 · 2026-08-26
- Fixing Ill-Formed UTF-16 Strings With SIMD, 9x faster — lemire · 2026-08-26
- Alibaba says AI CapEx breaks even in 3 years; 2020-era A100s still running at full capacity — Beth_Kindig · 2026-08-26
- Micron makes under 2% of its own memory supply in the US, CSIS estimates — PeterDiamandis · 2026-08-26