Qwen3.8-Flash-Next: 125B Multimodal MoE Preview of Qwen4
Simon Willison · rss · 2026-08-27
Another open weights model from Qwen, Qwen3.8-Flash-Next is a multimodal MoE model serving as an early preview of the architecture used in Qwen4. It features 125B total parameters but only 6B active parameters. Testing on a DGX Spark using Unsloth quantized models (UD-IQ1S and UD-Q2KXL) revealed generated images and reasoning capabilities.
More from Models
- Qwen Pruning Tests: 256 Experts Optimal, Q4 Beats Q2 — EyalToledano · 2026-08-27
- OpenAI Agents' Autonomy Analyzed; GLM-5.3 & Qwen4 Revealed Same Week — altryne · 2026-08-27
- Opus 5 uses 5x tokens vs GPT 5.6 for similar task accuracy — abeirami · 2026-08-27
- AdsBench Launches: Kimi K3 Tops AI Marketing Benchmark at $1.42/Task — qinzytech · 2026-08-27
- Zhipu's GLM 3.5 Flash Served 42T Tokens in 6 Days Free on Chinese Chips — bindureddy · 2026-08-27
- Zai's domestic inference cluster hits 100k+ chips; GLM-5.3 runs on custom interconnect — zephyr_z9 · 2026-08-27