Alibaba Unveils Qwen3.8-Flash Preview: MoE Model with Next-Gen Architecture
Alibaba_Qwen · x · 2026-08-26
Alibaba released Qwen3.8-Flash, a multimodal MoE model serving as an early preview of the Qwen4 architecture. With 125B total parameters but only 6B active per token, it offers unmatched cost-efficiency. The new architecture features GDN+QSA hybrid attention, N-gram Embedding, and the Muon optimizer, reducing training costs to 1/9th of Qwen3.7-Plus. It shows strong performance on benchmarks like DeepSWE, SWE-bench Pro, and MathVision. The production version will be available via QwenCloud API at $0.16/1M input tokens and $0.47/1M output tokens.
Related event: Alibaba unveils Qwen3.8-Flash preview with big cost cuts(2 posts)→
More from Infra
- CPU vs GPU vs TPU vs NPU vs LPU: 5 Hardware Architectures Explained — HankYeomans · 2026-08-26
- India holds just 0.2% of global compute capacity — himanshustwts · 2026-08-26
- Linux Foundation Celebrates 35th Anniversary, Open Source Targets AI — 0xsachi · 2026-08-26
- Open Source Facefusion Android App Runs Video Face Swap on Snapdragon NPU — Few_Caregiver8134 · 2026-08-26
- Apple M5 Mac Studio page features LM Studio for local AI — mattturck · 2026-08-26
- Chinese chip packaging firms invest $2.2B in expansion fueled by AI boom — pstAsiatech · 2026-08-26