GPT-6 Astra reportedly solves competition math without verbalized reasoning, with sharply worse monitorability
MattGarciaEth · x · 2026-09-04
Ryan Greenblatt highlights benchmark results suggesting GPT-6 Astra is a massive jump in opaque reasoning: it appears to solve hard competition math problems entirely "in its head," without verbalized reasoning, while prior models handled only basic word problems. He flags benchmark caveats — especially contamination — as a key concern.
UK AISI reportedly found Astra has much worse monitorability, making internal reasoning harder to audit. Greenblatt guesses the jump comes from architectural changes with increased serial depth, though a normal large-scale pretrain scale-up is plausible; similar jumps in future generations would be deeply concerning for oversight.
More from Models
- A brief history of modern open source AI, from a Hugging Face early investor — Borthwick · 2026-09-04
- Hugging Face claims 18 million people use its open-weights models — flowersslop · 2026-09-04
- Google AI Pro at $5/mo vs GPT Plus vs OpenCode Go: a coder's comparison — Old-Dish-7104 · 2026-09-04
- Community ranks 109 AI models by adjusted SEAL scores from Scale Labs — Hrstar1 · 2026-09-04
- GPT-6 Astra Sets Epoch ECI Record but Matches GPT-5.6 on AA Index, Sparking Benchmark Doubts — scaling01 · 2026-09-04
- GPT-6-Astra hands-on: smarter and creatively persistent, cuts 33k lines of SQL to 1.8k — i_dg23 · 2026-09-04