Qwen 3.8 Flash-Next: 6B-Active Open Model Beats Claude Opus 4.6 Max
kimmonismus · x · 2026-08-26
Alibaba's Qwen team released Qwen3.8-Flash-Next, an extremely sparse MoE model. While it has a total of 125B parameters (plus 51B n-gram embeddings), only 6B parameters are active per token. The model performs excellently on multiple benchmarks, beating Claude Opus 4.6 Max on 8 comparable metrics, including SWE-bench Pro (62.5), SWE-bench Multilingual, and CoWorkBench. It also supports native 256K context extendable to 1M, with significantly faster inference speeds at long context (7.6× faster prefill, 4.9× faster decode).
More from Models
- Tip: opencode's new model is reportedly a nerfed multimodal GLM-5.3 — PawelHuryn · 2026-08-26
- unsloth releases Qwen3.8-Flash-Next-FP8 at 186GB — LegacyRemaster · 2026-08-26
- Model Becomes Fastest Trending in Hugging Face History in 40 Minutes — MaziyarPanahi · 2026-08-26
- Qwen 3.8 Flash beats Opus on SWE-bench Pro — Hesamation · 2026-08-26
- Qwen3.8-Flash-Next: New Architecture Targets Ultimate Cost-Efficiency — tosh · 2026-08-26
- Minecraft clone fully vibecoded with local Qwen3.8-27b — liright · 2026-08-26