Inside KV-Cache Sharing: How DeepSeek CED and GLM 5.2 Differ

Developers break down KV-cache sharing designs: YOCO lets all layers in the model's second half share one KV cache; DeepSeek's CED applies a similar YOCO-style global branch per CSA2 layer, caching only encoder KV; GLM 5.2 scores per token while sharing block indices for flexibility and performance.

2026-09-11 ~ 2026-09-11 · 4 related posts

Full story(3 episodes)→