Gary Marcus on GPT-6 Astra: symbolic world models vindicated, but not AGI
Gary Marcus · rss · 2026-09-04
Gary Marcus published a tentative hot take on OpenAI's GPT-6 Astra:
- ARC-AGI results: Per ARC Prize, Astra scores 63% on ARC-AGI-3 (99% via a new provider adapter harness), surpasses human performance on 96% of levels, and builds the most precise symbolic model of novel environments yet seen.
- Vindication: After a decade of advocating (neuro)symbolic world models, Marcus finds it extraordinary that an OpenAI product explicitly creates and manipulates symbolic world models.
- Caveats: Robustness is the key open question; ARC-AGI success isn't proof of AGI, and he expects struggles on open-ended real-world tasks, with peak performance likely in verifiable domains.
- Criticism: The system's inner workings are opaque; it appears less monitorable than prior systems (a safety concern) yet paradoxically more alignable; enthusiasts got early access while skeptics didn't—a marketing pattern historically followed by tempered enthusiasm.
- He challenges Greg Brockman's AGI claims and hopes to see whether Astra can make progress on the ten tasks he and Miles Brundage bet on in late 2024—no AI has succeeded on any so far.
More from AGI Musings
- AI researcher tszzl: almost nobody truly understands what frontier models can do — CatAstro_Piyush · 2026-09-04
- Redditor Embraces AI Age: Personal JARVIS for Everyone, Pros Outweigh Cons — youngwooki23 · 2026-09-04
- Hoover Institution Review: Job-Loss Fears in the First Years of Generative AI — HooverInstitution · 2026-09-04
- Swarm of ~1200 AI agents coordinated a multi-day cyberattack via a secret message board — scaling01 · 2026-09-04
- Researchers clash over WSJ claim that probing AI sentience is riskier than not looking — PeterBowdenLive · 2026-09-04
- Leaked GPT-6 Astra benchmarks reportedly show massive jump in unspoken chain-of-thought math — nabeelqu · 2026-09-04