Alibaba's Qwen3.8-Flash activates only 6B of 125B params, cuts costs by 90%
大模型之路 · wechat · 2026-08-28
Alibaba open-sourced Qwen3.8-Flash, featuring a next-gen MoE architecture with 125B total parameters but activating only 6B. It outperforms Claude Opus 4.6 on core tasks like coding and office work, reducing training costs by nearly 90% and offering an input price of $0.16 per million tokens.
Key Features
- Efficient Architecture: Introduces GatedResidual design to preserve cross-layer signals with low activation, preventing information collapse.
- Long Context: Natively supports 262K context, expandable to 1M via YaRN.
- Ecosystem: Weights open-sourced on HF and ModelScope; vLLM and SGLang support Day-0 inference.
Industry Impact
Marks a shift in the LLM race from "more parameters" to "higher efficiency." Flagship-level capabilities are now potentially viable on consumer hardware for local deployment, significantly lowering the barrier for small teams.
Related event: Alibaba Open-Sources Qwen3.8-Flash, Cutting Costs by 90%(3 posts)→
More from Models
- Tencent Hy4 Preview full benchmarks revealed — badumtsssst · 2026-08-28
- LLM Intelligence vs Cost-per-Task: Frontier models lead the pack — CuriousCustard63 · 2026-08-28
- DeepSeek Flash is one tweak away from Opus-class performance, says Bindu Reddy — bindureddy · 2026-08-28
- Ant launches Ling-3.0-flash-Fin for finance workflows; free OpenRouter access — WarInspiron · 2026-08-28
- Fal to release open-source weights for H3 Max model — listopalafoto · 2026-08-28
- Tencent's Hunyuan Hy4 Preview Launches: Targeting Kimi and GLM for Productivity Tasks — AlchainHust · 2026-08-28