Alibaba's Qwen3.8-Flash activates only 6B of 125B params, cuts costs by 90%

大模型之路 · wechat · 2026-08-28

Alibaba open-sourced Qwen3.8-Flash, featuring a next-gen MoE architecture with 125B total parameters but activating only 6B. It outperforms Claude Opus 4.6 on core tasks like coding and office work, reducing training costs by nearly 90% and offering an input price of $0.16 per million tokens.

Key Features

Industry Impact

Marks a shift in the LLM race from "more parameters" to "higher efficiency." Flagship-level capabilities are now potentially viable on consumer hardware for local deployment, significantly lowering the barrier for small teams.

Related event: Alibaba Open-Sources Qwen3.8-Flash, Cutting Costs by 90%(3 posts)→

Original post →

More from Models

Models channel →