Claude Code auto-compact clarified: context is replaced by a summary, not kept in full

YashasGunderia · x · 2026-10-10

Lydia Hallie clears up a common misconception about Claude Code's auto-compact: when it runs, the whole conversation is replaced by a short summary — the last 1M tokens aren't kept around. The 967K figure only applies to 1M-context models.

Compacting is a single request that reads about as many tokens as the message sent right before it, and it's mostly cache reads if the cache is still warm. A replier confirms the same from their own usage. Useful practical detail for cost and context management in long-context agent sessions.

Related event: Anthropic Clarifies Claude Code Auto-Compact: Default 967K Threshold Burns Through Quota(4 posts)→

Original post →