Shared KV Cache in Looped Transformers Saves Memory and Boosts Performance

A new preprint highlighted by Cornell professor Yoav Artzi shows that sharing a single KV cache across recursions in looped Transformers saves 76-79% of context memory while achieving even lower perplexity.

2026-10-07 ~ 2026-10-07 · 2 related posts