Alibaba Unveils Qwen3.8-Flash Preview: MoE Model with Next-Gen Architecture

Alibaba_Qwen · x · 2026-08-26

Alibaba released Qwen3.8-Flash, a multimodal MoE model serving as an early preview of the Qwen4 architecture. With 125B total parameters but only 6B active per token, it offers unmatched cost-efficiency. The new architecture features GDN+QSA hybrid attention, N-gram Embedding, and the Muon optimizer, reducing training costs to 1/9th of Qwen3.7-Plus. It shows strong performance on benchmarks like DeepSWE, SWE-bench Pro, and MathVision. The production version will be available via QwenCloud API at $0.16/1M input tokens and $0.47/1M output tokens.

Related event: Alibaba unveils Qwen3.8-Flash preview with big cost cuts(2 posts)→

Original post →

More from Infra

Infra channel →