Qwen3.8-Flash-Next: GDN+QSA Hybrid Attention Cuts VRAM by 23.5GB

ying11231 · x · 2026-08-27

Qwen3.8-Flash-Next (Qwen4 arch preview) is a 125B MoE (6B active) with Day-0 support from SGLang. Architecture: Uses a hybrid of 36 Gated DeltaNet + 12 Qwen Sparse Attention layers for efficient long-context compute and precise retrieval. Utilizes 51B N-gram embeddings offloaded to host memory, saving 23.5 GiB VRAM (TP4) and boosting KV capacity by 78.5%. Performance: Achieves 540 tok/s decode on NVIDIA B200; HyperConnection Kernels deliver 2.05x kernel speedup and 7.6% end-to-end boost. Collaboration: Optimized with Alibaba Qwen, NVIDIA, AMD, and RadixArk, offering an NVFP4 quantized checkpoint.

Related event: Alibaba Open-Sources Qwen3.8-Flash-Next, an Ultra-Sparse MoE Preview of Qwen4(16 posts)→

Original post →

More from Infra

Infra channel →