Discussion: Does Claude Keep Chain-of-Thought Tokens in History?
Sufficient_Fox_4402 · reddit · 2026-07-20
Developers using models with deep reasoning capabilities (like Opus) raised questions about context token consumption: during multi-turn conversations, does the model retain the "thinking process" from previous messages in its context?
- Background: Early reasoning models typically discarded chain-of-thought tokens after a single response to save context space. However, modern intensive reasoning models can generate 5k-10k tokens of thought per turn. If counted towards the context, this severely impacts effective context management.
- Question: Do Claude Code or the Web version currently still discard these tokens after thinking is complete?
More from Models
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22
- Kimi K3 tops Gemini 3.6 Flash on four shared public benchmarks — ChrisGPT · 2026-07-22
- Google’s year-long pause in new base-model pretraining draws sharp criticism — teortaxesTex · 2026-07-22
- Current setup is 8,192 input tokens and 2,048 output tokens, with 8k/512 next — TheZachMueller · 2026-07-22
- Kimi K3 feels slower than K2.7, but stronger on long coding jobs and refactoring — Far-Presence2711 · 2026-07-22