Qwen3.8-Flash-Next architecture: 125B params, sparse MoE, 262K context

Alibaba_Qwen · x · 2026-08-26

vLLM announces support for Qwen3.8-Flash-Next, an ultra-sparse MoE multimodal model. It features 125B total parameters (including a 51B N-gram table) with only 6B active per token and native 262K context support (extendable to 1M via YaRN). The architecture combines Gated DeltaNet for history compression, Qwen Sparse Attention for precise retrieval, and MTP for speculative decoding. The N-gram table can be offloaded to host memory.

Related event: Alibaba Releases Qwen3.8-Flash-Next, an Ultra-Sparse Multimodal MoE Previewing Qwen4(10 posts)→

Original post →

More from Infra

Infra channel →