MLA-based KV cache costs 12 GB per million tokens, with KDA state at 230 MB BF16
zephyr_z9 · x · 2026-07-28
A post breaks down memory costs for an architecture using MLA layers and KDA state at 1M-token context.
It claims:
- KV cache from the MLA layers is about 12 GB per million tokens
- that is roughly 3× larger than the “whale” baseline
- the fixed KDA state is about 230 MB in BF16
The exchange is about whether the KV estimate is right, and whether BF16 is the right precision assumption.
More from Infra
- fmgo: call Apple's on-device Foundation Models from Go with no CGO and no Swift — Super_Run_8466 · 2026-09-23
- Huawei unveils Peerium architecture: nested BSP unifies million processors into one computer — Dr_Singularity · 2026-09-23
- Grok explains why DeepSeek picked DualPipe + ZeRO-1 over ZeRO-3 on 2048 H800s — TheZachMueller · 2026-09-23
- AI costs fall 47% per quarter, 4x faster than DNA sequencing: Epoch AI — daveholtz · 2026-09-23
- M5 Ultra LLM test: 4x faster prompt processing, but double the power draw — DigitalguyCH · 2026-09-23
- $500 of Dell OptiPlexes become a diskless netboot lab where AI agents can't brick the hardware — colinmcnamara · 2026-09-23