ARC-AGI Insight: More Reasoning Tokens Lead to Wrong Hypotheses; GPT-5.6 Max is Cheaper

GregKamradt · x · 2026-08-08

Greg Kamradt highlighted an analysis from ARC Prize explaining the counter-intuitive "zigzag" cost curve in reasoning models. Research reveals that higher reasoning performance often comes at a lower cost.

ARC officials explain that models running in Max reasoning mode actually consume fewer tokens overall than High mode. The reasons include:

For instance, on one ARC-AGI-2 task, High mode spent 98k tokens building a complex but wrong theory, whereas Max mode solved it correctly with only 76k tokens. This dynamic also underscores the outstanding cost-to-performance ratio demonstrated by DeepSeek.

Related event: DeepSeek V4 Flash Tops ARC-AGI Cost-Performance Frontier(2 posts)→

Original post →

More from Models

Models channel →