Leaked Qwen next-gen model shows record n-gram params, first two layers SWA-only
stochasticchasm · x · 2026-09-11
A technical breakdown of an apparent Qwen next-gen model architecture notes a gigantic n-gram parameter size — the largest seen yet, compared to 51B on qwen-3.8-flash-next, though this model is larger overall.
The author also flags that the first two of 20 layers use sliding-window attention only, speculating it's either a compute-saving design or analogous to first-k-dense patterns in MoEs.
Related event: Leaked Architecture Hints at Massive n-gram Head in Suspected Qwen Model(2 posts)→
More from Models
- DeepSeek v4.1 Flash Spotted Online, Authenticity Unverified — petrusenko_max · 2026-09-11
- Persimmon team members share months-in-the-making launch, research preview open — niloofar_mire · 2026-09-11
- Power user: Astra's usage limits are 'a joke' compared to Google's plan — MickeySteamboat · 2026-09-11
- GPT Astra Takes on Dominions 6, a Brutally Complex 4X Strategy Game — garden_frog · 2026-09-11
- Benchwarmer Tool Rebuilds Misleading AI Benchmark Charts and Recomputes the Winners — aronchick · 2026-09-11
- GPT-6 Astra autonomously flies a drone to find and follow a person, tops Drone-Bench — TheMoonMidas · 2026-09-11