Research: Inference-time Recurrence Cuts LLM Perplexity by 23% Without Weights Update

heghbalz · x · 2026-08-20

Introduces an inference-time recurrence mechanism allowing off-the-shelf, frozen LLMs to act as dynamical systems. By leaking deep representations back to shallow layers, it achieves significant performance gains without touching model weights: 23% reduction in perplexity and a 21% relative accuracy increase on GSM8k.

Original post →

More from Research

Research channel →