Qwen 3.8 Flash Next: Using Next-n-grams for efficient local inference
antirez · x · 2026-08-30
Antirez analyzed the architecture of Qwen 3.8 Flash Next, which uses a 51B table of n-grams at layer 2 to enrich token representations with a gating mechanism. This design allows masking and loading 16 small vectors from the SSD during layer 1 processing. Observers note that with the large n-grams table residing on SSD, Qwen 3.8 Flash Next with 2-bit quantization could be an optimal choice for local inference on 64GB MacBooks.
More from Infra
- DLSS 5 Video Player tested: Runs slow on 3090 Ti due to FP8 lack — fallengt · 2026-08-30
- Data Centers Drive US Reindustrialization: Benefits from Taxes to Jobs — GavinSBaker · 2026-08-30
- Running Qwen3.8-Next-Flash on 96GB RAM: offload the n-gram table to SSD — Iory1998 · 2026-08-30
- Vietnam's OneNexus quantizes GLM-5.3 to MXFP4 — xiaosun86 · 2026-08-30
- Pi Agent setup: Using Qwen 27B as an Oracle for co-development — Thrumpwart · 2026-08-30
- Full-stack AI EDA to disrupt chip design economics — ai · 2026-08-30