Claude engineer clarifies auto-compact: full summary replaces context, cache reads still burn quota
lydiahallie · x · 2026-10-10
Anthropic engineer Lydia Hallie clears up misconceptions about Claude Code's auto-compact:
- Auto-compact does summarize: when triggered, the whole conversation is replaced by a short summary — it never keeps the last 1M tokens around.
- The 967K threshold only applies to 1M-context models. Compacting is a single request reading roughly as many tokens as the message sent right before it, mostly cheap cache reads if the cache is warm.
- But waiting until 967K means every earlier message carries the full conversation history. Even with mostly cache hits, it still counts toward usage and adds up — which explains users burning 2%–15% of their 5h limits on a single auto-compact.
- Practical fix: /autocompact 400k triggers summarization at 400K so messages never balloon that big.
More from coding & agent
- A daily prompt that makes AI refactor your logs and errors for better debugging — DanielLockyer · 2026-10-10
- Microsoft pairs Surface Laptop Ultra with agent sandbox MXC and AI routing in Windows — pstAsiatech · 2026-10-10
- Open-source creative harness Voyager adds 8 languages, drives Blender, Resolve and 100+ tools — ravisparikh · 2026-10-10
- Claude Managed Agents finds 21 real bugs in 18k lines of code in 5 minutes for $8.36 — daniel_mac8 · 2026-10-10
- Testing an AI agent to book a hotel failed — and it still required manually entering credit card info — NielsRogge · 2026-10-10
- AI now writes 52% of merged code, but DX researcher says stop treating AICP as a KPI — rseroter · 2026-10-10