MLA-based KV cache costs 12 GB per million tokens, with KDA state at 230 MB BF16
zephyr_z9 · x · 2026-07-28
A post breaks down memory costs for an architecture using MLA layers and KDA state at 1M-token context.
It claims:
- KV cache from the MLA layers is about 12 GB per million tokens
- that is roughly 3× larger than the “whale” baseline
- the fixed KDA state is about 230 MB in BF16
The exchange is about whether the KV estimate is right, and whether BF16 is the right precision assumption.
More from Infra
- Compute, not algorithms, is the real moat in frontier AI — GavinSBaker · 2026-07-28
- Renting GPUs and open-weight models cut one AI bill from $1.2M to $100K — kimmonismus · 2026-07-28
- Running Kimi-K3 Locally Costs Up to 1 Million Euros, Jokes German Tech Blogger — FlorianGallwitz · 2026-07-28
- Bittensor updates $TAO emissions with a 61% gate that favors top-valued subnets — markjeffrey · 2026-07-28
- Qualcomm says NPU latency beats GPU and CPU for on-device AI agents — qdrant_engine · 2026-07-28
- A 70B GGUF model stalls on an AMD R9700 as VRAM hits 30 GB but RAM stays flat — Developer-Y · 2026-07-28