Prime Intellect compresses MLA KV cache in NVFP4, fitting ~50% more tokens than FP8
TheZachMueller · x · 2026-10-03
Prime Intellect shared its Decode NVFP4 KV compression approach: while TP4 has the lowest inter-token latency, every rank holds a full KV copy. Storing the MLA latent in NVFP4 (576 down to 352 bytes/row) fits 50% more cached tokens per decoder versus FP8. A native sparse-MLA kernel unpacks it on-chip and is being contributed to FlashInfer as an experimental operation.
More from Infra
- INT21: 2 engineers direct AI to build 20 inference engines in 2 weeks, up to 2.4× faster than SGLang — bingxu_ · 2026-10-03
- Traversal's 5 Levels of Self-Driving Production: Why Coding Agents Make Ops Harder — AI Engineer · 2026-10-03
- How DatologyAI Generated 12 Trillion Synthetic Tokens — And Fixed 4 Pipeline Bottlenecks — AI Engineer · 2026-10-03
- Price-Insensitive Buyer With Billions Seeks 500MW-2GW of Powered Data Center Land — JohnnyNi13 · 2026-10-03
- Running 256k-context open models on 2x RTX 3090 for months: a home server LLM retrospective — knighty1981 · 2026-10-03
- Runware launches Serverless GPUs: $0 while idle, from $0.63/GPU-hour — aziz4ai · 2026-10-03