22,022 API calls analyzed: Fable 5.1 burns ~31% more tokens, plus 3 quota traps
tenequm · reddit · 2026-09-16
Based on 22,022 of their own Claude Code API calls, a Reddit user breaks down three reasons Claude quota drains faster this week:
- Never cleanly communicated: the cache pricing discount only applies to API usage — Boris Cherny's price-cut announcement explicitly listed "Enterprise, API, and SDK customers"; subscriptions are not included.
- The announced one: the temporary +50% weekly boost became a permanent +25% from Sep 13, i.e. 17% less quota.
- Open bug: messages sent to finished subagents miss the cache entirely, so their full context gets re-billed on every request — examples show only 16K of 243K/399K written tokens read from cache.
Their companion analysis found Fable 5.1 consumes 31% more tokens per prompt than 5.0.
More from coding & agent
- Diorama gives OpenAI Codex coding agents a visual office you can watch work in real time — davidfromkansas · 2026-09-17
- Code-first, UI on top: building bespoke brand design tools with AI — floguo · 2026-09-17
- Study of 7 models across Claude Code, Codex, Pi: harness barely affects success but swings cost — DavideCrapis · 2026-09-17
- AI trading bot built with Jev is down 85%, owner shrugs it off — generativist · 2026-09-17
- Redditor's 3-Day SoL-Pi Test: Memory Objects Save ~12k Tokens Per Tool Run — Garblyx · 2026-09-17
- Reviewing AI code through Steve Jobs' lens: unseen internals deserve beauty too — sergeykarayev · 2026-09-17