Straight-up FP4 KV cache with no degradation: inference engineers take note

stochasticchasm · x · 2026-09-11

Commentary on a new model architecture: the KV cache is stored directly in FP4 with no observable degradation, and KVs are compressed via CSA — with a new "pure CSA2" design replacing the odd alternating CSA/HCA dual-frequency setup of v4. Combined, the model barely uses any KV cache, which the author says will make inference engineers' lives interesting; open questions include how kernel reduction order affects performance.

Related event: Inference-First Architecture Sparks Debate: FP4 KV Cache and Pure CSA2 Compression in Focus(11 posts)→

Original post →

More from Infra

Infra channel →