Claude Code quota drain traced to 967K default compact threshold; set /autocompact 400k to save
dotey · x · 2026-10-11
- Anthropic engineer Lydia Hallie clarified that Claude Code's auto-compact does summarize history, but on 1M-context models it waits until 967K tokens before triggering — so every turn before compaction carries a massive context base that quietly eats your 5-hour quota.
- Compaction itself isn't costly: the summary mostly runs on warm cache reads, and you only pay one small Cache Write for the short summary, in exchange for a much lower per-turn base afterward.
- Fix: run /autocompact 400k or set autoCompactWindow in /.claude/settings. to force earlier compaction.
- Quoting yetone: there's no universally optimal config; magpie's new "tune for you" feature replays compaction thresholds and cache durations locally against your real calls to find the most token-efficient settings per agent and model.
More from coding & agent
- AI Coding Bottleneck Shifts to Test Suites: Dev Wants 256GB RAM and 40 Cores Per Project — cjimti · 2026-10-11
- Reddit: Google pulls Gemini 3.1 Pro from Antigravity IDE, leaving devs on Flash — Adventurous-Week-399 · 2026-10-11
- Mine your agent session logs: 491 sessions reviewed for patterns — Teknium · 2026-10-11
- Nous Research consolidates Hermes Agent official plugins into a single mono-repo — Teknium · 2026-10-11
- Test suites become the bottleneck: dev ditches M5 Pro for beefy remote sandboxes — cjimti · 2026-10-11
- laya-guard: local guardrail for coding agents blocks 25/27 AgentDojo attacks at ~90ms — Extreme_Ad5709 · 2026-10-11