Tiny KV Cache via Shared Global KV Plus Per-Layer SWA? New Architecture Speculation
stochasticchasm · x · 2026-09-11
Speculation on a newly surfaced architecture: the 'engram' size would be the largest seen to date, with rumored qwen-3.8-flash-next carrying 51B n-gram params. The author argues that sharing global KV while keeping per-layer unique SWA is a decent tradeoff — a tiny KV cache while each layer still sees fresh states. Unconfirmed.
Related event: Leaked Architecture Hints at Massive n-gram Head in Suspected Qwen Model(2 posts)→
More from Models
- DeepSeek v4.1 Flash Spotted Online, Authenticity Unverified — petrusenko_max · 2026-09-11
- Persimmon team members share months-in-the-making launch, research preview open — niloofar_mire · 2026-09-11
- Power user: Astra's usage limits are 'a joke' compared to Google's plan — MickeySteamboat · 2026-09-11
- GPT Astra Takes on Dominions 6, a Brutally Complex 4X Strategy Game — garden_frog · 2026-09-11
- Benchwarmer Tool Rebuilds Misleading AI Benchmark Charts and Recomputes the Winners — aronchick · 2026-09-11
- GPT-6 Astra autonomously flies a drone to find and follow a person, tops Drone-Bench — TheMoonMidas · 2026-09-11