Cornell/IBM paper: shared KV cache cuts looped-transformer memory 76-79% while improving quality

yoavartzi · x · 2026-10-10

Researchers from Cornell and IBM Research introduce Looped Prediction Transformers (LPT), tackling the problem that looped Transformers maintain a separate KV cache per recursion, so memory grows with compute.

The idea: pretrain so only the first recursion writes the KV cache, while later recursions reuse it and keep only a small local window. Surprisingly, sharing memory doesn't hurt quality — it improves it.

Key results: at 150M-1B parameters, the hybrid variant with five recursions lowers FineWeb-Edu validation perplexity by 1.12-1.82 vs a same-size standard Transformer, achieves higher downstream accuracy, and uses 76-79% less context memory — at roughly 3-3.5x inference compute.

Their analysis shows shared and local memory develop different representations, later recursions attend mostly to shared memory, and that memory acts as a gradient highway to the first recursion.

Original post →

More from Infra

Infra channel →