Training-Free Memory Augmentation Makes Chain-of-Thought Cheaper and Faster

A new arXiv paper proposes training-free memory augmentation that moves reusable reasoning into the prompt, letting LLMs generate shorter chains of thought. The method gains 21.4 points on GSM8K and speeds up inference by nearly 1.5x.

2026-09-01 ~ 2026-09-01 · 2 related posts