Recirculation: Training-Free Inference-Time Architecture Boost for LLMs

The paper proposes an inference-time architecture enhancement technique called "Recirculation," which lets off-the-shelf foundation models improve their reasoning without retraining. Current findings: on Gemma3 it achieves a 23% reduction in perplexity and a 21% accuracy gain on GSM8k, and it consistently lowers perplexity across five model families. The highlight is zero training cost, making it worth watching whether it can become a general-purpose, low-cost way to enhance open-source models.

Confirmed

Why it matters

2026-08-19 ~ 2026-08-20 · 6 related posts

Primary sources