AA Intelligence Index under fire as Muse Spark 1.3 outranks GPT-6 Astra
_weiping · x · 2026-09-04
DanDr1s questions the credibility of the Artificial Analysis Intelligence Index after it ranked Muse Spark 1.3 (62) above GPT-6 Astra (61) — despite GPT-6 Astra scoring 98.6% on ARC-AGI-3, 97.6% on FrontierMath, and 100% on ExploitBench.
weiping adds that any single-number "benchmark" is inherently misleading, and benchmaxxing the index is worse than gaming any specific benchmark: it obscures key capabilities and leads people to dismiss meaningful benchmarks not included in the index.
More from Models
- ChatGPT's Reddit-flavored reply to a sexual assault victim sparks training-data backlash — airkatakana · 2026-09-04
- Astra Shows Near-Ideal Test-Time Scaling Gains on LifeSciBench Benchmark — soleio · 2026-09-04
- Astra reportedly hits 97% on ARC-AGI-3 without chain-of-thought — FeeAvailable3770 · 2026-09-04
- GPT-6 'Astra' Smashes ARC-AGI Records: 62.7% on Standard Harness, 99.9% on ARC-AGI-3 — repligate · 2026-09-04
- Browser QA harness: GLM 5.3 Flash beats DeepSeek V4 Flash Vision on screenshots — Certain_Pension6305 · 2026-09-04
- Reddit user flags Artificial Analysis as unreliable: same model shows conflicting scores — metigue · 2026-09-04