Compaction tosses 262GB of KV cache to keep 20KB, says Muennighoff
Muennighoff · x · 2026-10-11
Presenting at the MIT NLP Seminar on infinite test-time scaling and prefix sliding, Muennighoff estimates: with 32 dense layers, 4096 hidden dim, bf16, 500K context and a 5K compaction window, the KV cache is 262GB while compaction retains only 20KB — an enormous information loss.
He thinks compaction can already get us to models working autonomously for weeks on hard tasks, but it's very inefficient; replacing compaction may be the "final boss."
- KV cache: 32 × 4096 × 2 bytes × 500K × 2 ≈ 262GB
- Compaction output: 5000 × 4 bytes = 20KB
Related event: Muennighoff presents infinite test-time scaling at MIT(3 posts)→
More from Infra
- NVIDIA inference expert on when to keep, repurpose or replace aging GPUs — kimmonismus · 2026-10-11
- Independent Researcher Makes TPU Pallas top-k Bitwise Correct and 1.67x Faster — Francis_YAO_ · 2026-10-11
- AI agent tunes Triton kernels on AMD MI210, flipping grid order yields 1.29x speedup — zmkzmkz · 2026-10-11
- Engineer with $120k of GPUs: bought 75% before the price surge, don't follow me — TheZachMueller · 2026-10-11
- Migrating AI workloads often means 70%+ redevelopment, warns cloud analyst — DavidLinthicum · 2026-10-11
- Qwen3.8-Flash-Next on a 5090 beats Claude Code on airbench: 11 min vs 14 min — dh7net · 2026-10-11