Cornell/IBM paper: shared KV cache cuts looped-transformer memory 76-79% while improving quality
yoavartzi · x · 2026-10-10
Researchers from Cornell and IBM Research introduce Looped Prediction Transformers (LPT), tackling the problem that looped Transformers maintain a separate KV cache per recursion, so memory grows with compute.
The idea: pretrain so only the first recursion writes the KV cache, while later recursions reuse it and keep only a small local window. Surprisingly, sharing memory doesn't hurt quality — it improves it.
Key results: at 150M-1B parameters, the hybrid variant with five recursions lowers FineWeb-Edu validation perplexity by 1.12-1.82 vs a same-size standard Transformer, achieves higher downstream accuracy, and uses 76-79% less context memory — at roughly 3-3.5x inference compute.
Their analysis shows shared and local memory develop different representations, later recursions attend mostly to shared memory, and that memory acts as a gradient highway to the first recursion.
More from Infra
- When 2,000 agents share one CPU: Daytona on scaling isolated agent environments — mattturck · 2026-10-10
- Texas data center power queue hits 474GW, 90% from data centers — FinanceYF5 · 2026-10-10
- DDR5 hits $7,200 for 256GB as engineers treat RAM as an appreciating AI asset — CtrlAltDwayne · 2026-10-10
- Lemire reruns 2026 WebSocket benchmarks: his old Bun-vs-Node.js result was wrong, Anthropic bought Bun and Cloudflare bought Deno — lemire · 2026-10-10
- 4GB VRAM local LLM users: is there anything faster than llama.cpp? — your_real_Fathe_ · 2026-10-10
- Cascade GPU Topology Can Be Slower: The PCIe Hop Trap in Multi-GPU P2P — TheZachMueller · 2026-10-10