Charles Frye: KV Compression Fails Rarely but Expensively at Long Context, and New Models Need It Less
charles_irl · x · 2026-09-20
Weighing fp8e4m3 KV cache vs NVFP4, Charles Frye argues KV compression is "lowkey sus": failures (e.g. OOD inputs) are rare, but when they hit long-context, decode-heavy workloads the cost to users is high — and such failures are hard to benchmark. He adds that models from the last few months are much more KV-efficient, so compression gains matter less in the short-to-intermediate term.
Related event: KV Cache Compression Benefits Becoming Marginal, Says Frye(2 posts)→
More from Infra
- "There are more inference workloads in Heaven and Earth, Horatio" — a quip on overfit optimization — charles_irl · 2026-09-20
- Kimi subscriptions return after roughly two months, suggesting Moonshot found more compute — ChrisGPT · 2026-09-20
- HN: How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip — petrusenko_max · 2026-09-20
- FlashNorm: two lines of algebra buy 33-35% speedup — and a CUDA race that made the model echo the past — AI Engineer · 2026-09-20
- NEAR AI Brings Confidential Inference to Bittensor Subnet SayGm, an OpenRouter-Style Router — markjeffrey · 2026-09-20
- Game engines and inference engines both boil down to multi-user batching, and agentic bots — yunta_tsai · 2026-09-20