Reproducing Recirculation: Gemma 3 PPL Drops 23% Without Fine-tuning

CatAstro_Piyush · x · 2026-08-28

An independent researcher reproduced the Recirculation paper, confirming that injecting source layer 11 into destination layer 4 effectively reduces perplexity on the Gemma 3 1B model. The original paper proposes this as an inference-time architectural enhancement that boosts accuracy in generation and reasoning tasks without training. Tests on the Gemma3 family show a 23% reduction in PPL and a 21% accuracy increase on GSM8k.

Original post →

More from Research

Research channel →