Qwen 3.8 Flash-Next Released: 6B Sparse MoE Outperforms Claude Opus

SimplyAnnisa · x · 2026-08-26

Alibaba released Qwen 3.8 Flash-Next, a highly sparse MoE model with 6B active parameters. It features 125B total parameters and 51B additional n-gram embeddings, activating only 6B parameters per token.

The model beats Claude Opus 4.6 Max in 8 out of 9 comparable benchmarks. Key scores include:

Related event: Alibaba Open-Sources Qwen3.8-Flash-Next: 125B Ultra-Sparse MoE Previewing Qwen4 Architecture(11 posts)→

Original post →

More from Models

Models channel →