Straight-up FP4 KV cache with no degradation: inference engineers take note
stochasticchasm · x · 2026-09-11
Commentary on a new model architecture: the KV cache is stored directly in FP4 with no observable degradation, and KVs are compressed via CSA — with a new "pure CSA2" design replacing the odd alternating CSA/HCA dual-frequency setup of v4. Combined, the model barely uses any KV cache, which the author says will make inference engineers' lives interesting; open questions include how kernel reduction order affects performance.
More from Infra
- DeepSeek releases V4.1-Flash: 552B MoE claimed to beat V4-Pro on cost and speed — eyishazyer · 2026-09-11
- DOJ scrutinizes Nvidia's ~$20B Groq licensing deal over merger-review evasion — eyishazyer · 2026-09-11
- Persimmon Built on NVIDIA's 550B Nemotron 3 Ultra with Thousands of Blackwell GPUs — niloofar_mire · 2026-09-11
- NVIDIA details EPD disaggregation: up to 5x faster TTFT and 7x faster responses for multimodal serving — NVIDIAAI · 2026-09-11
- US and China Race to Build GPUs, UAE Builds Datacenters — Where's Europe? — tech__unicorn · 2026-09-11
- Longer Context = Faster Prefill? A Puzzling llama.cpp Benchmark Anomaly — Ekepa · 2026-09-11