Dev says AA index lost credibility the day Opus 5 outscored Fable 5 on paper
HarveenChadha · x · 2026-09-05
Harveen Chadha argues that eval scores are increasingly unrepresentative of real-world usage: he stopped trusting the Artificial Analysis index the day Opus 5 outscored Fable 5, and finds it worse still that Muse Spark 1.3 beats Astra on the leaderboard. His takeaway: rankings and lived experience are diverging — don't pick models by benchmark scores alone.
Related event: Hands-on tests call out Muse Spark 1.3 benchmark mismatch(4 posts)→
More from Models
- Yoav Goldberg: capabilities once dependent on the harness are now baked into the model — yoavgo · 2026-09-05
- Yoav Goldberg: Ark's harness was simply bad, and OpenAI's fix was obvious — yoavgo · 2026-09-05
- Anthropic Fable 5.1 vs OpenAI Astra: analyst teases a clear winner — dylan522p · 2026-09-05
- Reviewer: OpenAI's GPT-6-Astra finally 'gets what you mean,' with Fable-level intelligence and real gains in game dev — pvncher · 2026-09-05
- Meta ships Muse Spark 1.3 with max reasoning, pitching frontier performance at non-frontier prices — AIatMeta · 2026-09-05
- LLMs are now making up words that don't exist, not just jargon — StewartalsopIII · 2026-09-05