Leaked Architecture Hints at Massive n-gram Head in Suspected Qwen Model
Analysis of a leaked architecture, suspected to be a new Qwen model, reveals the largest known n-gram (multi-token prediction) head, with the first two layers using only SWA—possibly trading globally shared KV plus independent SWA for a very small KV cache.
2026-09-11 ~ 2026-09-11 · 2 related posts
- Leaked Qwen next-gen model shows record n-gram params, first two layers SWA-only — stochasticchasm · 2026-09-11
- Tiny KV Cache via Shared Global KV Plus Per-Layer SWA? New Architecture Speculation — stochasticchasm · 2026-09-11