Tokenomics 101: understanding input, output, and cached token pricing in the AI era
Aizkmusic · x · 2026-09-02
The author links to their 'Tokenomics 101' blog post (following up on the eye-watering token bill from their Fable 5.1 test). Key points:
- The world's appetite for tokens has exploded: thousands of context tokens were standard years ago, while a million is now typical even for open source models. Agents can run for hours or days and spawn subagents — the post cites Jarred Sumner's Rust rewrite of Bun, which at peak ran four Claude Code workflows with sixteen Claudes each (64 Claudes at once).
- Virtually all major providers price tokens three ways: input, output, and cached. Cached tokens hit a server that recently processed the same prompt, saving enormous compute and time; the discrete turn-by-turn nature of LLM conversations makes them naturally well suited to caching, which is mostly handled automatically behind the scenes today.
More from Models
- Fable 5.1 beats Opus 5 in price/performance on AA Index — JasonBotterill · 2026-09-02
- WSJ: Gemini 3.8 Flash drops tomorrow, preferred over Opus in internal coding tests — kimmonismus · 2026-09-02
- Fable 5.1 solves reading comprehension perfectly with zero reasoning — Sauers_ · 2026-09-02
- Fable 5.1 drains Claude 5-hour limit in under 30 minutes, users report — robleclerc · 2026-09-02
- Replit Announces Atlas: A Versatile Autoregressive Multimodal Model — gowthami_s · 2026-09-02
- Astra Model Touted as Impressive by Industry Observers — inductionheads · 2026-09-02