Recirculation adds inference-time memory to LLMs without retraining, boosting Gemma3 accuracy by 21%

pbaylies · x · 2026-08-20

The paper introduces "Recirculation," an inference-time architectural enhancement that gives pretrained Transformers a form of working memory by feeding deep-layer information back into earlier layers.

Key Mechanics:

Results on Gemma3:

The findings suggest that memory mechanisms may already exist within pretrained models, and recirculation simply provides a way to persist them.

Related event: Recirculation: Training-Free Architectural Enhancement for LLM Reasoning(7 posts)→

Original post →

More from Research

Research channel →