Artificial Analysis isn't broken: self-funded benchmarks, $13k spent on one model
Antblue · reddit · 2026-09-11
Responding to recent complaints that Artificial Analysis is "broken" or "bought out", the author counters point by point: the Intelligence Index is a weighted aggregate of 10 evaluations, most with published arXiv papers; only AA-Briefcase is private, and the aggregation methodology is public. AA runs independent benchmarks on its own funding without ads—it spent $13,129 testing Fable 5.1 alone, and benchmarks every new model.
Deepseek V4.1-Flash shows why aggregate scores miss the picture: the 552B model scores the same (40) as the 180B Qwen 3.8-Flash-Next, yet matches or exceeds it on most individual evals, beats GPT-6 Astra (Max) on AutomationBench-AA (agentic SaaS workflows), but falls behind on AA-Omniscience Non-Hallucination Rate, where open-weight models usually lead. Strengths and weaknesses should be celebrated.
Advice: read the individual evaluations, the papers, and the aggregation method before complaining. The author is unaffiliated with AA.
More from Models
- OpenAI reportedly pauses new $200 Pro subs as GPT-6 Astra demand strains capacity — WhatTheLJW · 2026-09-11
- DeepSeek-V4.1-Flash hits Ollama: 552B MoE backbone with 1M context via KV cache compression — ollama · 2026-09-11
- Multi-agent evals still undecided, but colocated async RL training is catching on — stochasticchasm · 2026-09-11
- Does DeepSeek V4.1-Flash's SWA Bounded Replay sacrifice recall to save KV cache memory? — Top-Handle-5728 · 2026-09-11
- ChatGPT starts inserting ads after each answer, users complain — mansithole6 · 2026-09-11
- Reddit users mourn old coding flow: new models spend 10 minutes overthinking and miss the point — snoosnoosewsew · 2026-09-11