Alibaba Qwen Releases Qwen3.8-Flash: 6B Active Params, 9x Cost Reduction

千问大模型 · wechat · 2026-08-26

Alibaba Qwen released the Qwen3.8-Flash multimodal MoE model with a 125B main size but only 6B active parameters per token. It features a GDN+QSA hybrid architecture, GatedResidual multi-branch streams, and N-gram Embedding for capacity expansion. It natively supports 262k context (extendable to 1M). Training costs are reduced to 1/9 of Qwen3.7-Plus while improving coding and office tasks. Priced at 1 RMB/M input tokens and 3 RMB/M output tokens. Weights for the new architecture Qwen3.8-Flash-Next are open-sourced.

Related event: Alibaba Open-Sources Qwen3.8-Flash-Next: 125B Ultra-Sparse MoE Previewing Qwen4 Architecture(11 posts)→

Original post →

More from Models

Models channel →