SGLang ships day-0 recipes for Qwen3.8-Flash: 176B params, 6B active, verified on RTX PRO 6000 and DGX Spark
NVIDIAAI · x · 2026-09-08
Alibaba's Qwen team launched Qwen3.8-Flash (a preview of the Qwen4 architecture), and SGLang shipped day-0 deployment recipes, each verified on real hardware:
- Architecture: 176B parameters with 51B of N-gram embeddings that scale capacity at almost no extra per-token compute (6B activated per token); the embedding table sits in host memory with async prefetch instead of GPU memory
- Attention: GDN + QSA hybrid attention balances memory efficiency and precise retrieval on long-horizon tasks; Gated Residual gives 4 lanes for inter-layer information flow; trained with Muon
- Verified recipes: 1x RTX PRO 6000 (PLE table in system RAM), 1x DGX Spark (local NVMe), 2x DGX Spark with TP=2, plus NVFP4 support via RadixArk and NVIDIA ModelOpt exports
- Caveat: model support isn't in a tagged release yet — build from PR #36497; the team has fixed several issues reported by the community
More from Infra
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11
- Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device — lee_stott · 2026-09-11