Qwen3.8-Flash-Next architecture: 125B params, sparse MoE, 262K context
Alibaba_Qwen · x · 2026-08-26
vLLM announces support for Qwen3.8-Flash-Next, an ultra-sparse MoE multimodal model. It features 125B total parameters (including a 51B N-gram table) with only 6B active per token and native 262K context support (extendable to 1M via YaRN). The architecture combines Gated DeltaNet for history compression, Qwen Sparse Attention for precise retrieval, and MTP for speculative decoding. The N-gram table can be offloaded to host memory.
More from Infra
- Alibaba's Qwen3.8 with 51B N-gram embeddings now available on SGLang — Alibaba_Qwen · 2026-08-26
- Edge sensor nodes struggle to disconnect from hyper-centralized foundation models — curious_vii · 2026-08-26
- Teaser: Apple MLX King to Reveal Local Model Advances — dscape · 2026-08-26
- Glean saves 81% on token costs with right-sized intelligence strategy — _akhaliq · 2026-08-26
- MiniMax H3 on Mac: 480p/5s video in under 8 minutes with Turbo LoRA ComfyUI workflow — xNightWardenx · 2026-08-26
- Domyn Uses NVIDIA Nemotron to Build 263B Model for Regulated Industries — NVIDIA Developer · 2026-08-26