Training-free memory augmentation makes CoT cheaper: +21-29 pts and 1.5x speedup

dreamwieber · x · 2026-09-01

An arXiv paper, Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning, offers a new way to cut CoT costs: move reusable reasoning from generation into the prompt.

Key ideas:

Results: Memory improves prompt-based Chain-of-Draft compression by 21.4 / 28.0 / 29.5 / 6.61 points on GSM8K, MATH, BBH, and MMLU-Sci, while achieving 1.14-1.49x latency speedup over standard CoT. It's compatible with token-level, trace-level, and inference-state compression.

Related event: Training-Free Memory Augmentation Makes Chain-of-Thought Cheaper and Faster(2 posts)→

Original post →

More from coding & agent

coding & agent channel →