Qwen 3.8 Flash Next rumored: 125B + 51B N-gram lookup tables, launches tomorrow

teortaxesTex · x · 2026-08-25

Rumors say Qwen 3.8 Flash Next (125B params + 51B N-gram embeddings, 6B active, based on next-gen Qwen 4 architecture) launches tomorrow. superelesha analyzes the 51B N-gram tables: a massive trainable lookup-memory that hashes recent token N-grams and mixes fetched vectors into the hidden state, so the main model spends less effort on frequent local patterns.

Memory math (ideal 4bpw):

Caveat: real Q4 is usually heavier, and some quantizers keep embedding tables in FP16/BF16 — the dtype of the 51B parameters is the key question of this release.

Related event: Alibaba Teases Qwen3.8-Flash-Next Open Release as Preview of Qwen4 Architecture(8 posts)→

Original post →

More from Infra

Infra channel →