ARC-AGI Insight: More Reasoning Tokens Lead to Wrong Hypotheses; GPT-5.6 Max is Cheaper
GregKamradt · x · 2026-08-08
Greg Kamradt highlighted an analysis from ARC Prize explaining the counter-intuitive "zigzag" cost curve in reasoning models. Research reveals that higher reasoning performance often comes at a lower cost.
ARC officials explain that models running in Max reasoning mode actually consume fewer tokens overall than High mode. The reasons include:
- High mode often burns through tokens elaborating on an incorrect hypothesis, eventually failing to produce a usable prediction.
- Max mode correctly combines fundamental rules (like color conservation) earlier, hitting the right answer without unnecessary exploration.
For instance, on one ARC-AGI-2 task, High mode spent 98k tokens building a complex but wrong theory, whereas Max mode solved it correctly with only 76k tokens. This dynamic also underscores the outstanding cost-to-performance ratio demonstrated by DeepSeek.
Related event: DeepSeek V4 Flash Tops ARC-AGI Cost-Performance Frontier(2 posts)→
More from Models
- Qwen3.8-Max beats Gemini 3.5 Flash by 8.5 points in tests — usamawahabkhan · 2026-08-08
- Semianalysis Deep Dive on Gemini 3.5 Pro: Performance and Architecture Insights — Charuru · 2026-08-08
- DeepSeek V4 Flash Appears on ARC Prize Leaderboard — tosh · 2026-08-08
- Pokee AI Launches Isaac Model with 10M-Token Context and API — Kyrannio · 2026-08-08
- Opinion: Google or Meta Could Win the AI Race by Dropping a Better Open-Source Model Than K3 — bindureddy · 2026-08-08
- Why LLMs Can't Count Tokens: The Need for Intermediate Steps — ctjlewis · 2026-08-08