Training-Free Memory Augmentation Makes Chain-of-Thought Cheaper and Faster
A new arXiv paper proposes training-free memory augmentation that moves reusable reasoning into the prompt, letting LLMs generate shorter chains of thought. The method gains 21.4 points on GSM8K and speeds up inference by nearly 1.5x.
2026-09-01 ~ 2026-09-01 · 2 related posts
- Training-free memory augmentation makes CoT cheaper: +21-29 pts and 1.5x speedup — dreamwieber · 2026-09-01
- Training-free memory augmentation makes compressed CoT faster: +21.4 pts on GSM8K — rohanpaul_ai · 2026-09-01