Real Task Cost Across GPT, Claude, Gemini, Kimi: 10.6x Spread Despite Only 2x Price Difference, Hidden Reasoning Tokens Blamed
pixelo2323 · reddit · 2026-07-23
Benchmarking 10 realistic product tasks against live APIs of OpenAI, Anthropic, Gemini, and Kimi reveals a 10.6x total cost spread despite only 2x published rate difference. The culprit: reasoning and thinking tokens billed at output rate but hidden from responses. One-word classification answer cost 197 invisible reasoning tokens. Related to CostBench (ACL 2026) and TerminalWorld findings. Full methodology and results on GitHub.
Related event: Benchmark Reveals 10x Cost Gap Among Top AI Models(2 posts)→
More from Models
- OpenAI opens GPT-5.6 to the public across ChatGPT, Codex and API — emmanuelvivier · 2026-07-23
- Laguna at low quant seems to overthink and burn through context fast — IUseClifford · 2026-07-23
- Qwen-Image-3.0 gets put through layout-heavy tests against GPT-Image-2 — Scobleizer · 2026-07-23
- Dean Ball says Kimi is strong at coding, but open-weight economics still hurt — koltregaskes · 2026-07-23
- Simon Willison says loops are becoming obsolete as models handle long tasks natively — teropa · 2026-07-23
- A Qwen 3.6 27B derivative claims to cut thinking tokens by more than 90% — AppealSame4367 · 2026-07-23