Reddit benchmark finds a 10.6x real-task cost spread across GPT, Claude, Gemini and Kimi
pixelo2323 · reddit · 2026-07-23
A Reddit post compares real task costs across GPT, Claude, Gemini, and Kimi using 10 realistic product tasks such as classification, RAG QA, multi-turn conversation, and agentic plan-and-execute flows.
The main finding is a 10.6x cost spread even though the public price cards differ by only 2x. The poster says the gap is driven largely by hidden reasoning and thinking tokens, which are billed at output rates but not shown in the response. One example: a one-word classification answer consumed 197 invisible reasoning tokens.
The post also connects the result to related research, including CostBench, which finds models often fail to pick cost-optimal plans, and TerminalWorld, which reports that failed agent attempts burn disproportionately more tokens than successful ones. Methodology, raw results, and prompts are public on GitHub.
Related event: Benchmark Reveals 10x Cost Gap Among Top AI Models(2 posts)→
More from Models
- Five frontier models all solved the same bugs, but cost varied 14x and Claude refused 40% — PromptPhanter · 2026-07-23
- Antirez says Laguna S2.1 cannot write correct Italian via the official API — antirez · 2026-07-23
- Model lineage matters more than API traces, says Eyisha Zyer — eyishazyer · 2026-07-23
- Claude users warned usage limits may reset if Opus 5 launches today — CtrlAltDwayne · 2026-07-23
- Reddit screenshots suggest Laguna says Poolside in English, Qwen in Chinese — Serious-Affect-6410 · 2026-07-23
- X rumor says GPT-5.6, Cerebras 750 token/s release and Claude Opus 5 may land today — Scobleizer · 2026-07-23