Shared memory makes looped transformers better: 76-79% less KV memory and lower perplexity

yoavartzi · x · 2026-10-07

Cornell professor Yoav Artzi shares his team's new preprint, "The Surprising Effectiveness of Shared Memory in Looped Transformers" (arXiv:2610.02383).

Key idea: Looped transformers reuse the same layers per token, but each recursion writes its own KV cache, so memory grows with compute. The authors instead pretrain models to share memory: only the first recursion writes the cache, and later recursions read it while keeping a short local window.

Results: At 150M–1B parameters, sharing memory improves rather than hurts quality. With five recursions, the hybrid variant lowers validation perplexity on FineWeb-Edu by 1.12–1.82 vs. a same-size standard transformer while using 76–79% less context memory.

Why it works: Shared and local memory develop different representations; later recursions attend mostly to shared memory, which also acts as a gradient highway back to the first recursion, improving backprop.

Original post →

More from Research

Research channel →