GPT-6 Closes the Coding Gap but Trails on Deep Reasoning — Why Google May Reach AGI First
aditipawarr · reddit · 2026-09-04
Amid debate over GPT-6 Astra topping Artificial Analysis, the author argues it exposes a weakness: near-perfect long-horizon stability, but notably below Fable on Humanity's Last Exam, only matching 8-month-old Gemini 3.1 Pro with tools.
Key takes:
- Deep reasoning (HLE), DeepSWE and Arc-AGI together determine big AA jumps
- Astra is an assistant-style step for coding/long-horizon only; OpenAI staff hint a stronger model is coming this year
- Google, with superior data stack, semi-solved long-horizon via 3.8 Flash, and a unique world-model stack, is the author's AGI favorite
- Anthropic expected to stay laser-focused on narrow domains
More from AGI Musings
- Where's My Mansion? The Flaw in the AI-Will-Make-Millions-Millionaires Story — AICopyLab · 2026-09-04
- Garrison Lovely announces 'Obsolete,' an AI-critical book endorsed by Nobel laureate Acemoğlu — GarrisonLovely · 2026-09-04
- Bratton amplifies view that halting AI would 'radically impoverish' human existence — bratton · 2026-09-04
- Is Astra AGI? Five contradictory answers that are all true at once — shaunralston · 2026-09-04
- 3-Month Reflection: LLMs Boosted My Productivity Less Than 50%, and Their Intelligence Is Nothing Like Ours — sebkrier · 2026-09-04
- Berlin Art Week forum 'New Conditions: Art After AI' explores how AI reshapes art — matdryhurst · 2026-09-04