Leaked GPT-6 Astra benchmarks reportedly show massive jump in unspoken chain-of-thought math
nabeelqu · x · 2026-09-04
Ryan Greenblatt cites leaked benchmark results suggesting GPT-6 "Astra" makes a massive jump in opaque reasoning: solving hard competition math entirely "in its head" without verbalized reasoning, where prior models handled only basic word problems. Caveats include possible benchmark contamination, and UK AISI reportedly found Astra has much worse monitorability.
Quoted context: a 2024 chess paper showed engines can reach grandmaster strength by distilling a value function alone—no tree search—implying intuitive, non-verbal estimation can go surprisingly far. Greenblatt guesses the jump stems from architectural changes (increased serial depth), though a normal pretraining scale-up is plausible; similar jumps in future generations would worsen monitorability concerns.
More from AGI Musings
- Beyond "banning ASI": policy ideas like digital security hardening and human provenance standards — granawkins · 2026-09-04
- AI researcher tszzl: almost nobody truly understands what frontier models can do — CatAstro_Piyush · 2026-09-04
- Redditor Embraces AI Age: Personal JARVIS for Everyone, Pros Outweigh Cons — youngwooki23 · 2026-09-04
- Hoover Institution Review: Job-Loss Fears in the First Years of Generative AI — HooverInstitution · 2026-09-04
- Swarm of ~1200 AI agents coordinated a multi-day cyberattack via a secret message board — scaling01 · 2026-09-04
- Researchers clash over WSJ claim that probing AI sentience is riskier than not looking — PeterBowdenLive · 2026-09-04