Qwen3.8-Flash-Next: GDN+QSA Hybrid Attention Cuts VRAM by 23.5GB
ying11231 · x · 2026-08-27
Qwen3.8-Flash-Next (Qwen4 arch preview) is a 125B MoE (6B active) with Day-0 support from SGLang. Architecture: Uses a hybrid of 36 Gated DeltaNet + 12 Qwen Sparse Attention layers for efficient long-context compute and precise retrieval. Utilizes 51B N-gram embeddings offloaded to host memory, saving 23.5 GiB VRAM (TP4) and boosting KV capacity by 78.5%. Performance: Achieves 540 tok/s decode on NVIDIA B200; HyperConnection Kernels deliver 2.05x kernel speedup and 7.6% end-to-end boost. Collaboration: Optimized with Alibaba Qwen, NVIDIA, AMD, and RadixArk, offering an NVFP4 quantized checkpoint.
More from Infra
- RootCrak builds x402 security layer for autonomous agent transactions — Thionne_WTZ · 2026-08-27
- Max Hodak: Anonymous model testing routed data to Chinese datacenter — ohlennart · 2026-08-27
- Opinion: Why Targeting Data Centers is an Environmentalist Mistake — AndyMasley · 2026-08-27
- Advocating for Independent Secure Clusters: Open Science Needs Open Compute — gajesh · 2026-08-27
- Self-hosting LLMs on Budget Hardware: Principles, Optimization, and Benchmarks — jflesch · 2026-08-27
- Stas Bekman's ML Engineering open book gets massive update, now 498 pages — StasBekman · 2026-08-27