Real Task Cost Across GPT, Claude, Gemini, Kimi: 10.6x Spread Despite Only 2x Price Difference, Hidden Reasoning Tokens Blamed

pixelo2323 · reddit · 2026-07-23

Benchmarking 10 realistic product tasks against live APIs of OpenAI, Anthropic, Gemini, and Kimi reveals a 10.6x total cost spread despite only 2x published rate difference. The culprit: reasoning and thinking tokens billed at output rate but hidden from responses. One-word classification answer cost 197 invisible reasoning tokens. Related to CostBench (ACL 2026) and TerminalWorld findings. Full methodology and results on GitHub.

Related event: Benchmark Reveals 10x Cost Gap Among Top AI Models(2 posts)→

Original post →

More from Models

Models channel →