Sharing one KV cache across recursions improves looped Transformers, preprint finds
yoavartzi · x · 2026-10-07
A new preprint argues against training looped Transformers with a separate KV cache per recursion: sharing a single cache across recursions is a net-positive inductive bias, improving performance and stability while cutting memory at the same FLOPs. Researchers including Yoav Artzi amplified the finding.
Related event: Shared KV Cache in Looped Transformers Saves Memory and Boosts Performance(2 posts)→
More from Research
- eigenrobot: automating math papers is easy, and most of economics and theoretical physics is next — eigenrobot · 2026-10-07
- François Fleuret: math is unique in that its truths are "context free" — francoisfleuret · 2026-10-07
- Two Years After First Reasoning Model, AI Has Produced '20 Fields Medals' of New Math — __nmca__ · 2026-10-07
- NUS releases SafeActBench: 656 cases reveal where tool-using agents break the evidence-to-action chain — NationalUniversityofSingapore · 2026-10-07
- MEND: RL for flow models via proximal velocity matching beats Flow-GRPO in 100 vs ~4k updates — UTEXAS · 2026-10-07
- JLD: perceptual distance from a frozen encoder's Jacobian, fitted in 35s from 100 images, beats LPIPS and DISTS — Shreshth Saini · 2026-10-07