Qwen Team Analyzes Qwen3.8-Next Architecture Design
Qwen · hf · 2026-09-01
The Qwen team released a post on the design of the Qwen3.8-Next architecture. Qwen3.8-Flash-Next uses a sparse mixture-of-experts architecture, combining hybrid gated delta-net and sparse attention layers, gated residual branches, and off-accelerator n-gram embeddings to improve efficiency, capability, and training stability.
More from Models
- Meme Suggests AI Models Are Distilling Each Other — toomanynamesaretook · 2026-09-01
- Hypothesis on Opus 5 leakiness: Mixed old and new training formats — Ratter · 2026-09-01
- Opus training format shift: From plain text to XML tags — Ratter · 2026-09-01
- Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored Trends on HF — DavidAU · 2026-09-01
- Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory — Runjia Qian · 2026-09-01
- Zhipu Releases INT4 & MXFP4 Versions of GLM-5.3 Flash — HaihaoShen · 2026-09-01