Prime Intellect compresses MLA KV cache in NVFP4, fitting ~50% more tokens than FP8

TheZachMueller · x · 2026-10-03

Prime Intellect shared its Decode NVFP4 KV compression approach: while TP4 has the lowest inter-token latency, every rank holds a full KV copy. Storing the MLA latent in NVFP4 (576 down to 352 bytes/row) fits 50% more cached tokens per decoder versus FP8. A native sparse-MLA kernel unpacks it on-chip and is being contributed to FlashInfer as an experimental operation.

Original post →

More from Infra

Infra channel →