GPT-6 Astra uses reasoning tokens in only 27% of ARC-AGI-3 actions, hinting at a leap in opaque reasoning
mhmazur · x · 2026-09-05
New data point from ARC-AGI-3 on GPT-6 Astra
Under low reasoning effort in a standard harness:
- GPT-5.6 Sol used reasoning tokens in 98% of its actions while playing ARC-AGI-3 public games
- GPT-6 Astra used them in only 27% of actions (!)
Semi-private games showed the same pattern, suggesting the difference isn't driven by public benchmark contamination and may instead reflect architectural changes in how the GPT models work.
In the surrounding discussion, the shift is framed as a potentially massive jump in opaque reasoning: GPT-6 Astra appears able to solve hard competition math problems entirely "in its head" without verbalized reasoning, where prior models could only handle basic word problems. The author flags this as concerning if the benchmark results are representative, while noting the highlighted caveats. Still third-party analysis pending further validation.
Related event: GPT-6 Astra Beats Sol max 2x on ARC v3 at Low Reasoning(2 posts)→
More from Models
- First hands-on: GPT-6 Astra nails Blender modeling of a prison phone in one pass — AIandDesign · 2026-09-05
- GPT-6 Astra appears in Codex during testing, availability scope unclear — dotey · 2026-09-05
- GPT-6 Astra spotted in early tests, reportedly beating Fable benchmark — Angaisb_ · 2026-09-05
- User flags ChatGPT refusing to share official election website URLs — mimi10v3 · 2026-09-05
- Anthropic model formalizes Fermat's Last Theorem in Lean, 13.4M lines closing 100-theorem benchmark — littmath · 2026-09-05
- Astra nails a game-ready prison phone model in Blender on first try — AIandDesign · 2026-09-05