MLA-based KV cache costs 12 GB per million tokens, with KDA state at 230 MB BF16

zephyr_z9 · x · 2026-07-28

A post breaks down memory costs for an architecture using MLA layers and KDA state at 1M-token context.

It claims:

The exchange is about whether the KV estimate is right, and whether BF16 is the right precision assumption.

Original post →

More from Infra

Infra channel →