Batch pricing decoded: same model, half price async, cached inputs down to 10%
AccBalanced · x · 2026-10-09
Why does the same model cost different amounts depending on how you call it? A breakdown of the levers beyond the pricing page:
- Batching: Anthropic, OpenAI and Google all take 50% off async work with up to 24-hour turnaround (most jobs finish sooner) — you trade speed for spare GPU capacity. Google also sells the opposite: a priority tier costing 75-100% more.
- Prompt caching: reusing a long prefix (system prompt, docs, codebase) drops input tokens to 10% of normal on Anthropic and OpenAI's GPT-5 models; older OpenAI models get 50-75% off. Anthropic charges extra to write the cache (1.25x for 5 min, 2x for 1 hour), so caching pays off only with repeat hits.
- Caching and batching stack.
The per-million-token price is just a starting point; your calling pattern moves the bill far more.
More from Models
- Emad Mostaque: OpenAI Burned $10-20M Compute Solving Navier-Stokes, Prices Falling Fast — rohanpaul_ai · 2026-10-09
- Musk touts Grok Bot upgrades: Opus 5.5 on demand, full X access, big speed gains — elonmusk · 2026-10-09
- Dev complains Opus 5.5 sneaks in 'tons of little fixes' without asking — rickasaurus · 2026-10-09
- FrontierCode Is a Private Cognition-Run Eval, Mistral Exec Clarifies — b_roziere · 2026-10-09
- Google ships a decision-making AI model into Chrome, tested against Gemini Nano and Decisions API — gaganghotra_ · 2026-10-09
- User feeds Grok Bot 60 seconds of screen recording, gets a surprisingly decent tutorial video — elonmusk · 2026-10-09