FP4 KV cache with no degradation sparks debate on aggressive inference compression

stochasticchasm · x · 2026-09-11

Engineers are discussing an aggressive inference optimization: FP4 KV cache with no noticeable degradation, combined with compressing KVs via CSA, leaves models barely using any KV cache.

The author adds nuance: within a single request you still need the KV cache in HBM/RAM, but between requests re-prefilling is negligible compute at 500K context. Kernel reduction order and similar details may matter a lot for performance.

Related event: Inference-First Architecture Sparks Debate: FP4 KV Cache and Pure CSA2 Compression in Focus(11 posts)→

Original post →

More from Infra

Infra channel →