Compaction tosses 262GB of KV cache to keep 20KB, says Muennighoff

Muennighoff · x · 2026-10-11

Presenting at the MIT NLP Seminar on infinite test-time scaling and prefix sliding, Muennighoff estimates: with 32 dense layers, 4096 hidden dim, bf16, 500K context and a 5K compaction window, the KV cache is 262GB while compaction retains only 20KB — an enormous information loss.

He thinks compaction can already get us to models working autonomously for weeks on hard tasks, but it's very inefficient; replacing compaction may be the "final boss."

Related event: Muennighoff presents infinite test-time scaling at MIT(3 posts)→

Original post →

More from Infra

Infra channel →