Claude Code auto-compact clarified: context is replaced by a summary, not kept in full
YashasGunderia · x · 2026-10-10
Lydia Hallie clears up a common misconception about Claude Code's auto-compact: when it runs, the whole conversation is replaced by a short summary — the last 1M tokens aren't kept around. The 967K figure only applies to 1M-context models.
Compacting is a single request that reads about as many tokens as the message sent right before it, and it's mostly cache reads if the cache is still warm. A replier confirms the same from their own usage. Useful practical detail for cost and context management in long-context agent sessions.