Qwen 3.8 Flash-Next Released: 6B Sparse MoE Outperforms Claude Opus
SimplyAnnisa · x · 2026-08-26
Alibaba released Qwen 3.8 Flash-Next, a highly sparse MoE model with 6B active parameters. It features 125B total parameters and 51B additional n-gram embeddings, activating only 6B parameters per token.
The model beats Claude Opus 4.6 Max in 8 out of 9 comparable benchmarks. Key scores include:
- SWE-bench Pro: 62.5
- SWE-bench Multilingual: 81.0
- CoworkBench: 73.9
- JobBench: 55.7
- Toolathlon: 73.5
- IFBench: 81.3
More from Models
- Zhipu GLM-5.3-Flash Offers 50% Discount for Two Weeks — Zai_org · 2026-08-26
- GLM-5.3-Flash adopts sparse + linear attention hybrid architecture — multimodalart · 2026-08-26
- GLM-5.3-Flash Uses Hybrid Attention to Cut Long-Context Costs — multimodalart · 2026-08-26
- Zhipu releases GLM-5.3-Flash: smaller size, Opus 4.8 level performance — airesearch12 · 2026-08-26
- Zhipu quietly releases glm-5.3-flash model — koltregaskes · 2026-08-26
- mLateOn beats massive Qwen models in performance — antoine_chaffin · 2026-08-26