Sharing one KV cache across recursions improves looped Transformers, preprint finds

yoavartzi · x · 2026-10-07

A new preprint argues against training looped Transformers with a separate KV cache per recursion: sharing a single cache across recursions is a net-positive inductive bias, improving performance and stability while cutting memory at the same FLOPs. Researchers including Yoav Artzi amplified the finding.

Related event: Shared KV Cache in Looped Transformers Saves Memory and Boosts Performance(2 posts)→

Original post →

More from Research

Research channel →