Alibaba releases Qwen3.8-Flash-Next: activates only 6B parameters per token with 50% price cut

智东西 · wechat · 2026-08-27

Alibaba's Qwen team open-sourced Qwen3.8-Flash-Next, a preview based on the next-generation Qwen4 architecture. With 125B total parameters, the MoE model activates only 6B parameters per token, reducing training costs to 1/9th of its predecessor. Key innovations include Qwen Sparse Attention (QSA), Gated Residual (GR), and N-gram Embedding. It supports a native 262k context window, extendable to 1 million tokens. It outperforms DeepSeek-V4-Flash and Claude-Opus-4.6 on multiple benchmarks. API pricing is set at 1 RMB input / 3 RMB output per million tokens, 33% lower than competitors.

Related event: Alibaba Releases Qwen3.8-Flash Multimodal MoE Model to Wide Acclaim(26 posts)→

Original post →

More from Models

Models channel →