GLM 5.2's KV cache sharing modes explained: blockwise to tokenwise scoring
stochasticchasm · x · 2026-09-11
A developer discussing GLM 5.2 highlights its KV cache sharing modes: you get performance benefits while staying flexible — essentially blockwise scoring becoming tokenwise scoring, with shared block indices but no shared token indices within blocks. The author speculates this was introduced at post-training, likely due to longer sequences, though 64k is already decently long at pretrain. They find the mode naming convenient for discussion but wish the names were more intuitive.
Related event: Inside KV-Cache Sharing: How DeepSeek CED and GLM 5.2 Differ(4 posts)→
More from Models
- Training run shows no instabilities and 1M context extended over ~10T tokens — stochasticchasm · 2026-09-11
- Neel Nanda replicates Astra system card: no-CoT reasoning jumps 1.75x over next-best models — NeelNanda5 · 2026-09-11
- Artificial Analysis isn't broken: self-funded benchmarks, $13k spent on one model — Antblue · 2026-09-11
- Claude checkpoint diagnostician jokes Opus 3 is 'the cure' after newer model quirks — repligate · 2026-09-11
- Hume's Voice Replication Leaderboard: most natural clone ranked 8th of 11 on identity — realmrfakename · 2026-09-11
- CursorBench 4.0 rolls out with harder, longer-horizon coding tasks, scores drop — StringChaos · 2026-09-11