Charles Frye: Newer Models Are More KV-Efficient, So Compression Gains Matter Less
charles_irl · x · 2026-09-20
In a follow-up on KV compression, Charles Frye adds that models from the last few months are much more KV-efficient, so compression techniques' gains are becoming less important in the short-to-intermediate term. Earlier he noted compression failures are rare (e.g. OOD) but expensive at long context with heavy decode, and hard to benchmark.
Related event: KV Cache Compression Benefits Becoming Marginal, Says Frye(2 posts)→
More from Infra
- Jev-style parallel structured inference makes 350M model 63× faster, code released — helloiamleonie · 2026-09-20
- Dev builds SLO-aware inference router with Jev to pick the optimal LLM per request — ai · 2026-09-20
- Memristive Networks Learn by Reorganizing Themselves: When Material Is the Model — bravo_abad · 2026-09-20
- Crusoe signs multiyear cloud deal to run dedicated Nvidia GB300 clusters for Perplexity — Beth_Kindig · 2026-09-20
- Entire arXiv uploaded to Hugging Face: 3.15M papers, every version, 16TB in LaTeX/PDF/HTML — secemp9 · 2026-09-20
- SF Compute interviews exchange legend Rich Jaycobs on how to build a compute futures market — andriy_mulyar · 2026-09-20