Artificial Analysis isn't broken: self-funded benchmarks, $13k spent on one model

Antblue · reddit · 2026-09-11

Responding to recent complaints that Artificial Analysis is "broken" or "bought out", the author counters point by point: the Intelligence Index is a weighted aggregate of 10 evaluations, most with published arXiv papers; only AA-Briefcase is private, and the aggregation methodology is public. AA runs independent benchmarks on its own funding without ads—it spent $13,129 testing Fable 5.1 alone, and benchmarks every new model.

Deepseek V4.1-Flash shows why aggregate scores miss the picture: the 552B model scores the same (40) as the 180B Qwen 3.8-Flash-Next, yet matches or exceeds it on most individual evals, beats GPT-6 Astra (Max) on AutomationBench-AA (agentic SaaS workflows), but falls behind on AA-Omniscience Non-Hallucination Rate, where open-weight models usually lead. Strengths and weaknesses should be celebrated.

Advice: read the individual evaluations, the papers, and the aggregation method before complaining. The author is unaffiliated with AA.

Original post →

More from Models

Models channel →