Why does GPT-6 Astra get mogged on the AA index? Caballero asks
ethanCaballero · x · 2026-09-04
Researcher Ethan Caballero questions GPT-6 Astra's weak showing on the Artificial Analysis index, a counterpoint to claims that it beats Fable 5.1 across nearly all benchmarks — suggesting results may vary depending on the evaluation source.
More from Models
- OpenAI's Astra can now layout and route PCBs, sparking hardware engineering debate — MikePFrank · 2026-09-04
- Ethan Mollick: Astra just takes action, spinning up agents on vague requests — emollick · 2026-09-04
- 'AGI is 74% deepswe': GPT-6-Astra benchmark results become an AI-circle meme — amaarora · 2026-09-04
- Grok 4.7 reportedly days away, trained on SpaceX engineering data; Grok 4.6 already ties GPT-6 Astra at 61 — XFreeze · 2026-09-04
- Ex-OpenAI safety lead Miles Brundage: if your primary emotion on AI isn't concern, you're misreading it — Miles_Brundage · 2026-09-04
- Gary Marcus on GPT-6 Astra: symbolic world models are vindication, but no proof of AGI — GaryMarcus · 2026-09-04