Alibaba releases Qwen3.8-Flash-Next: activates only 6B parameters per token with 50% price cut
智东西 · wechat · 2026-08-27
Alibaba's Qwen team open-sourced Qwen3.8-Flash-Next, a preview based on the next-generation Qwen4 architecture. With 125B total parameters, the MoE model activates only 6B parameters per token, reducing training costs to 1/9th of its predecessor. Key innovations include Qwen Sparse Attention (QSA), Gated Residual (GR), and N-gram Embedding. It supports a native 262k context window, extendable to 1 million tokens. It outperforms DeepSeek-V4-Flash and Claude-Opus-4.6 on multiple benchmarks. API pricing is set at 1 RMB input / 3 RMB output per million tokens, 33% lower than competitors.
Related event: Alibaba Releases Qwen3.8-Flash Multimodal MoE Model to Wide Acclaim(26 posts)→
More from Models
- ThursdAI Recap: GLM 5.3 and Qwen 27B Drop, OpenAI Pauses RL for Safety — altryne · 2026-08-28
- Qwen3.8-Flash-Next runs 162k context on dual 7900 XTX — BigYoSpeck · 2026-08-28
- Glitch Exposes Gemini's Internal Thoughts — Regular_Preference64 · 2026-08-28
- GLM-5.2 Introduces Monitors to Combat Reward Hacking in RL — burny_tech · 2026-08-28
- Anthropic Luna Max test shows generous limits, high speed — timpera · 2026-08-28
- User calls Grokbot 'terrible', cites missing tasks — krishnan · 2026-08-28