Claude Code quota drain explained: auto-compact waits until ~967K tokens by default
udmrzn · x · 2026-10-11
The community long suspected Claude Code wasn't compressing context at all, silently eating the 5-hour quota. Anthropic engineer Lydia Hallie clarified that auto-compact does summarize: once triggered, the entire history is replaced by a short summary, not stuffed with hidden tokens.
The real problem is the aggressive default: on 1M-context models, compaction only kicks in near 967K tokens. Until then, every turn carries a 800K+ token base, so even cache-hit reads rapidly burn the quota. Fears that compaction ruins caching are misplaced — the summary is mostly hot-cache reads, and you pay one small cache write for a much smaller ongoing base.
Practical fixes:
- Run /autocompact 400k or set autoCompactWindow in /.claude/settings. to force earlier compaction;
- Personal plans keep prompt cache alive 1 hour but subagents only 5 minutes — tune promptCacheTtl and subagentPromptCacheTtl;
- When the main session runs Opus, switch background subagents to lighter models to stretch the quota.
More from coding & agent
- Reddit: Google pulls Gemini 3.1 Pro from Antigravity IDE, leaving devs on Flash — Adventurous-Week-399 · 2026-10-11
- Mine your agent session logs: 491 sessions reviewed for patterns — Teknium · 2026-10-11
- Nous Research consolidates Hermes Agent official plugins into a single mono-repo — Teknium · 2026-10-11
- Test suites become the bottleneck: dev ditches M5 Pro for beefy remote sandboxes — cjimti · 2026-10-11
- laya-guard: local guardrail for coding agents blocks 25/27 AgentDojo attacks at ~90ms — Extreme_Ad5709 · 2026-10-11
- Sherlock game shows small-agent swarm plus strong reasoner solves 96/100 murder cases — No_Yogurtcloset_7050 · 2026-10-11