Leaked GPT-6 Astra benchmarks reportedly show massive jump in unspoken chain-of-thought math

nabeelqu · x · 2026-09-04

Ryan Greenblatt cites leaked benchmark results suggesting GPT-6 "Astra" makes a massive jump in opaque reasoning: solving hard competition math entirely "in its head" without verbalized reasoning, where prior models handled only basic word problems. Caveats include possible benchmark contamination, and UK AISI reportedly found Astra has much worse monitorability.

Quoted context: a 2024 chess paper showed engines can reach grandmaster strength by distilling a value function alone—no tree search—implying intuitive, non-verbal estimation can go surprisingly far. Greenblatt guesses the jump stems from architectural changes (increased serial depth), though a normal pretraining scale-up is plausible; similar jumps in future generations would worsen monitorability concerns.

Original post →

More from AGI Musings

AGI Musings channel →