Shared memory makes looped transformers better: 76-79% less KV memory and lower perplexity
yoavartzi · x · 2026-10-07
Cornell professor Yoav Artzi shares his team's new preprint, "The Surprising Effectiveness of Shared Memory in Looped Transformers" (arXiv:2610.02383).
Key idea: Looped transformers reuse the same layers per token, but each recursion writes its own KV cache, so memory grows with compute. The authors instead pretrain models to share memory: only the first recursion writes the cache, and later recursions read it while keeping a short local window.
Results: At 150M–1B parameters, sharing memory improves rather than hurts quality. With five recursions, the hybrid variant lowers validation perplexity on FineWeb-Edu by 1.12–1.82 vs. a same-size standard transformer while using 76–79% less context memory.
Why it works: Shared and local memory develop different representations; later recursions attend mostly to shared memory, which also acts as a gradient highway back to the first recursion, improving backprop.
More from Research
- COSMI composes single-object captures into 222k multi-object interaction sequences, 30x larger than prior sets — UniTuebingen · 2026-10-07
- Learned latent protein languages cut structure-prediction perplexity 34% and run ~1000x faster than AlphaFold2 — Mahdi Pourmirzaei · 2026-10-07
- UWaterloo's IGMBench tests world editing in Minecraft and Terraria; best agent solves 78.2% of tasks — UWaterloo · 2026-10-07
- Polar probe paper accepted at COLM 2026: linearly decoding semantic geometry in LLMs — JeanRemiKing · 2026-10-07
- AI math proofs shift to open-source-style collaboration, square packing shows — ctjlewis · 2026-10-07
- Overmind: Open Platform That Turns Production Traces Into Fine-Tuning Data for Agents — cneuralnetwork · 2026-10-07