Dev says AA index lost credibility the day Opus 5 outscored Fable 5 on paper

HarveenChadha · x · 2026-09-05

Harveen Chadha argues that eval scores are increasingly unrepresentative of real-world usage: he stopped trusting the Artificial Analysis index the day Opus 5 outscored Fable 5, and finds it worse still that Muse Spark 1.3 beats Astra on the leaderboard. His takeaway: rankings and lived experience are diverging — don't pick models by benchmark scores alone.

Related event: Hands-on tests call out Muse Spark 1.3 benchmark mismatch(4 posts)→

Original post →

More from Models

Models channel →