GLM 5.2's KV cache sharing modes explained: blockwise to tokenwise scoring

stochasticchasm · x · 2026-09-11

A developer discussing GLM 5.2 highlights its KV cache sharing modes: you get performance benefits while staying flexible — essentially blockwise scoring becoming tokenwise scoring, with shared block indices but no shared token indices within blocks. The author speculates this was introduced at post-training, likely due to longer sequences, though 64k is already decently long at pretrain. They find the mode naming convenient for discussion but wish the names were more intuitive.

Related event: Inside KV-Cache Sharing: How DeepSeek CED and GLM 5.2 Differ(4 posts)→

Original post →

More from Models

Models channel →