Alibaba Qwen Releases Qwen3.8-Flash: 6B Active Params, 9x Cost Reduction
千问大模型 · wechat · 2026-08-26
Alibaba Qwen released the Qwen3.8-Flash multimodal MoE model with a 125B main size but only 6B active parameters per token. It features a GDN+QSA hybrid architecture, GatedResidual multi-branch streams, and N-gram Embedding for capacity expansion. It natively supports 262k context (extendable to 1M). Training costs are reduced to 1/9 of Qwen3.7-Plus while improving coding and office tasks. Priced at 1 RMB/M input tokens and 3 RMB/M output tokens. Weights for the new architecture Qwen3.8-Flash-Next are open-sourced.
More from Models
- Zhipu GLM-5.3-Flash Offers 50% Discount for Two Weeks — Zai_org · 2026-08-26
- GLM-5.3-Flash Uses Hybrid Attention to Cut Long-Context Costs — multimodalart · 2026-08-26
- Zhipu releases GLM-5.3-Flash: smaller size, Opus 4.8 level performance — airesearch12 · 2026-08-26
- Zhipu quietly releases glm-5.3-flash model — koltregaskes · 2026-08-26
- mLateOn beats massive Qwen models in performance — antoine_chaffin · 2026-08-26
- Zai Launches 320B Parameter GLM-5.3-Flash Model — TheZachMueller · 2026-08-26