GPT-6 Astra uses reasoning tokens in only 27% of ARC-AGI-3 actions, hinting at a leap in opaque reasoning

mhmazur · x · 2026-09-05

New data point from ARC-AGI-3 on GPT-6 Astra

Under low reasoning effort in a standard harness:

Semi-private games showed the same pattern, suggesting the difference isn't driven by public benchmark contamination and may instead reflect architectural changes in how the GPT models work.

In the surrounding discussion, the shift is framed as a potentially massive jump in opaque reasoning: GPT-6 Astra appears able to solve hard competition math problems entirely "in its head" without verbalized reasoning, where prior models could only handle basic word problems. The author flags this as concerning if the benchmark results are representative, while noting the highlighted caveats. Still third-party analysis pending further validation.

Related event: GPT-6 Astra Beats Sol max 2x on ARC v3 at Low Reasoning(2 posts)→

Original post →

More from Models

Models channel →