DeepSeek cut per-token KV cache size by 54x in nine months

zephyr_z9 · x · 2026-09-10

DeepSeek has compressed per-token KV cache size by 54x over the past nine months, per a share on X — a striking inference-efficiency gain, since KV cache dominates long-context memory footprints.

Original post →

More from Infra

Infra channel →