Qwen 3.8 Flash Next: Using Next-n-grams for efficient local inference

antirez · x · 2026-08-30

Antirez analyzed the architecture of Qwen 3.8 Flash Next, which uses a 51B table of n-grams at layer 2 to enrich token representations with a gating mechanism. This design allows masking and loading 16 small vectors from the SSD during layer 1 processing. Observers note that with the large n-grams table residing on SSD, Qwen 3.8 Flash Next with 2-bit quantization could be an optimal choice for local inference on 64GB MacBooks.

Related event: antirez: Qwen 3.8 Flash Next's 51B N-gram Table Could Be Best Local Inference Option on 64GB Macs(7 posts)→

Original post →

More from Infra

Infra channel →