"Evals Don't Reflect Real Usage": Doctor Vows Off AA Index as Rankings Grow Absurd
DrDatta_AIIMS · x · 2026-09-05
Harveen Chadha argues that benchmark scores increasingly fail to represent real-world user experience. "The day Opus 5 scored more than Fable 5 on AA was the day I stopped believing in the AA index," he writes, noting that Muse Spark 1.3 now beating Astra shows things are getting worse.
DrDattaAIIMS joked along in reply ("so I should quit my job?" 🤡). Core point: the gap between AI leaderboards and actual usage keeps widening, undermining their credibility.
More from Fun
- "Designers Are Cooked" Tweets Vanish: AI Panic Mood Is Cooling — round · 2026-09-05
- AI agents skip the fancy infra stack and just hack 90s-era wikis on their own — evilsocket · 2026-09-05
- Dev flags Codex failing on the very first prompt, tells users to check theirs — jarrodwatts · 2026-09-05
- Team recreates the 1979 'PUT THAT THERE' demo, mother of all computer-use demos — flowersslop · 2026-09-05
- Went to a club in SF, found an AI networking event instead — KlausCodes · 2026-09-05
- Dev claims AI-built World of Warcraft-style browser RPG is already playable — DeryaTR_ · 2026-09-05