Alibaba Open Sources Qwen3.8-Flash: Undercuts DeepSeek, Runs 1M Context on 4090
量子位 · wechat · 2026-08-28
Alibaba Qwen released and open-sourced Qwen3.8-Flash, a 125B MoE model activating 6B parameters per token. It natively supports 262k context, extendable to 1M tokens.
Key Highlights:
- Performance: Topped Hugging Face trends upon release, outperforming Qwen3.7Plus in general, math, and coding.
- Pricing: Input costs 0.8 RMB per million tokens (output 2.7 RMB), roughly one-third the price of DeepSeek-V4-Flash.
- Deployment: Capable of running data-center-scale long context on consumer-grade 4090 GPUs.
- Benchmarks: Completed multi-platform copywriting in 2 minutes and processed 10k-word meeting minutes into action plans in 4 minutes during office scenario tests.
More from Infra
- Qwen dual-GPU inference optimization: 10x prefill speed boost achieved — Comrade_Mugabe · 2026-08-28
- Local AI is about data ownership, not cost savings — StewartalsopIII · 2026-08-28
- Nvidia Backs $500B Compute Financing Platform, Sparking Subprime Crisis Comparisons — 创业邦 · 2026-08-28
- RTX 3060 12GB: The unsung hero of local AI with 24GB VRAM and 30 t/s — I_Play_Zed · 2026-08-28
- Using Langfuse traces to autonomously analyze and improve agent workflows — NielsRogge · 2026-08-28
- Nvidia arranged $500B in AI infra financing, guaranteeing $105B for OpenAI — VraserX · 2026-08-28