Reddit benchmark finds a 10.6x real-task cost spread across GPT, Claude, Gemini and Kimi

pixelo2323 · reddit · 2026-07-23

A Reddit post compares real task costs across GPT, Claude, Gemini, and Kimi using 10 realistic product tasks such as classification, RAG QA, multi-turn conversation, and agentic plan-and-execute flows.

The main finding is a 10.6x cost spread even though the public price cards differ by only 2x. The poster says the gap is driven largely by hidden reasoning and thinking tokens, which are billed at output rates but not shown in the response. One example: a one-word classification answer consumed 197 invisible reasoning tokens.

The post also connects the result to related research, including CostBench, which finds models often fail to pick cost-optimal plans, and TerminalWorld, which reports that failed agent attempts burn disproportionately more tokens than successful ones. Methodology, raw results, and prompts are public on GitHub.

Related event: Benchmark Reveals 10x Cost Gap Among Top AI Models(2 posts)→

Original post →

More from Models

Models channel →