Tiny KV Cache via Shared Global KV Plus Per-Layer SWA? New Architecture Speculation

stochasticchasm · x · 2026-09-11

Speculation on a newly surfaced architecture: the 'engram' size would be the largest seen to date, with rumored qwen-3.8-flash-next carrying 51B n-gram params. The author argues that sharing global KV while keeping per-layer unique SWA is a decent tradeoff — a tiny KV cache while each layer still sees fresh states. Unconfirmed.

Related event: Leaked Architecture Hints at Massive n-gram Head in Suspected Qwen Model(2 posts)→

Original post →

More from Models

Models channel →