Astra tops every benchmark but stagnates on Arena, sparking distrust of Artificial Analysis index
brandon_galang · x · 2026-09-04
Quoting theo's post, brandongalang questions the credibility of the Artificial Analysis intelligence index after finding GPT-6 Astra ranked behind MUSE SPARK 1.3 MAX — despite Astra's dominant scores across reported benchmarks.
brandon adds that Astra seems static on Arena too, and says he no longer takes the Artificial Analysis index seriously. The open question: is the third-party index's sampling/weighting broken, or are the reported benchmarks overfit?
More from Models
- OpenAI's Astra can now layout and route PCBs, sparking hardware engineering debate — MikePFrank · 2026-09-04
- Ethan Mollick: Astra just takes action, spinning up agents on vague requests — emollick · 2026-09-04
- 'AGI is 74% deepswe': GPT-6-Astra benchmark results become an AI-circle meme — amaarora · 2026-09-04
- Grok 4.7 reportedly days away, trained on SpaceX engineering data; Grok 4.6 already ties GPT-6 Astra at 61 — XFreeze · 2026-09-04
- Ex-OpenAI safety lead Miles Brundage: if your primary emotion on AI isn't concern, you're misreading it — Miles_Brundage · 2026-09-04
- Gary Marcus on GPT-6 Astra: symbolic world models are vindication, but no proof of AGI — GaryMarcus · 2026-09-04