Leaked Qwen next-gen model shows record n-gram params, first two layers SWA-only

stochasticchasm · x · 2026-09-11

A technical breakdown of an apparent Qwen next-gen model architecture notes a gigantic n-gram parameter size — the largest seen yet, compared to 51B on qwen-3.8-flash-next, though this model is larger overall.

The author also flags that the first two of 20 layers use sliding-window attention only, speculating it's either a compute-saving design or analogous to first-k-dense patterns in MoEs.

Related event: Leaked Architecture Hints at Massive n-gram Head in Suspected Qwen Model(2 posts)→

Original post →

More from Models

Models channel →