Gemini bills thinking tokens on top of output while OpenAI counts them inside, dev's measurements show
qaiser_mehdi · reddit · 2026-09-07
A developer building a cost meter measured how providers count thinking tokens and found three providers with two opposite conventions:
- Google: verified across three gemini-3.6-flash calls, thoughtstokencount sits entirely outside candidatestokencount and is billed on top. In one call thinking was 589 of 750 tokens (79% of the call)—pricing by the 144 output tokens alone reported about a fifth of the real cost.
- OpenAI: reasoningtokens is a subset of completiontokens; adding them double counts.
- Anthropic: doesn't report thinking separately at all, folding it into outputtokens.
None of the three throws an error if you get it wrong. Caveat: one model, three calls on free-tier access. The verification script is open-sourced and reproducible with a free AI Studio key; the author asks whether other providers use Google's convention.
More from Models
- GPT-6 Astra high cut latency from 30s to 11s within hours — BLUECOW009 · 2026-09-07
- GPT-6 Astra builds a stunning Chicxulub impact animation in pure three.js — omarsar0 · 2026-09-07
- OpenAI's merettm says models could be far better at math, they just haven't focused on it — JeffLadish · 2026-09-07
- User prompts GPT-6 Astra to build its own Minecraft harness, record and edit a demo video — Angaisb_ · 2026-09-07
- Mathematician: latest models Sol 5.6 and Fable 5.1 can't solve any of my real problems — burny_tech · 2026-09-07
- User Praises OpenAI's Offer of Three Credit Resets — shekitup · 2026-09-07