No-Training Depth Boost? Recirculation Paper
teortaxesTex · x · 2026-08-19
A paper titled 'Recirculation' proposes an inference-time architectural enhancement that injects top-layer activations into bottom layers at the next step. This allows the model to act as a dynamical system to track belief states without retraining. It achieves a 23% perplexity reduction and a 21% accuracy boost on GSM8k for Gemma3, with minimal latency overhead during generation.
Related event: Recirculation Boosts LLM Reasoning Without Any Retraining(3 posts)→
More from Research
- Correcting Agentic Index: Introducing Cost-Per-Test and Pareto Frontiers Analysis — MikePFrank · 2026-08-20
- RL Math Part 14: Baselines, Advantage Function, and Actor-Critic — ShawnHymel · 2026-08-20
- GPT-5 and Gemini 2.5 Pro win gold medals at International Astronomy Olympiad — hhsun1 · 2026-08-20
- Relaxing chain rule yields new divergence; Tilted ERM improves generalization — burny_tech · 2026-08-20
- New paper: AI agent risks evolve from agency to autonomy to control — rohanpaul_ai · 2026-08-20
- Steerling-8B: Interpretable diffusion model trained with built-in explainability — burny_tech · 2026-08-20