ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium
i_dg23 · x · 2026-09-04
mhmazur breaks down the cost math on ARC-AGI-3's public game ft09: max reasoning used 10x the reasoning tokens of medium (41k vs 4k), but solved the game in 75 actions vs 91. Fewer actions meant less context replay, dropping max's per-game cost to $111 vs medium's $133. Across all semi-private games, this efficiency explains why max cost $26k overall versus medium's $48k — more thinking, less interacting, cheaper overall.
More from Models
- Claude Fable 5.1 Launches, Early Users Say It One-Shots the Best Websites of Any Model — repligate · 2026-09-04
- GPT-6 Astra debuts at No.1 on Terminal-Bench, 1.9% ahead of Claude Fable 5.1 — sandersted · 2026-09-04
- ARC-AGI-3 is now saturated, prompting calls for new benchmarks ASAP — kimmonismus · 2026-09-04
- OpenAI Researcher roon: GPT-6 Astra Will Be Obsolete in Weeks — Tolopono · 2026-09-04
- Reasoning effort switching without breaking cache is live in Codex and Claude — altryne · 2026-09-04
- antirez benchmarks DeepSeek v4 Flash vs GLM 5.3 Flash at Q2/Q4/mixed quants — antirez · 2026-09-04