"Evals Don't Reflect Real Usage": Doctor Vows Off AA Index as Rankings Grow Absurd

DrDatta_AIIMS · x · 2026-09-05

Harveen Chadha argues that benchmark scores increasingly fail to represent real-world user experience. "The day Opus 5 scored more than Fable 5 on AA was the day I stopped believing in the AA index," he writes, noting that Muse Spark 1.3 now beating Astra shows things are getting worse.

DrDattaAIIMS joked along in reply ("so I should quit my job?" 🤡). Core point: the gap between AI leaderboards and actual usage keeps widening, undermining their credibility.

Related event: Users Question Artificial Analysis Index as Real-World Testing Contradicts Rankings(5 posts)→

Original post →

More from Fun

Fun channel →