A benchmark joke says Anthropic doing badly is the fastest way to lose trust in it
teortaxesTex · x · 2026-07-25
A joking post says the easiest way to undermine Lisan’s trust in any benchmark is to show Anthropic doing badly on it.
The attached screenshot shows a discussion around Opus 5 ECI: one user claims it should be around 163.5 based on public benchmarks, says Anthropic’s internal AECl is 162.1, and notes that it is higher than Mythos 5. The punchline is the complaint that public benchmarks should be harder because the current ones are too easy to game or over-interpret.
Related event: Anthropic's Benchmark Scores Spark Community Trust Crisis(2 posts)→
More from Fun
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Someone built a website where you can sign up for AI not to kill you — motionbynick · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Meme: Engineers Unleash 10,000 Claude Sub-Agents on Friday Afternoon to Clear a Week's Work — _jaydeepkarale · 2026-09-11
- AI safety isn't a coordinated cabal: half the field has posted their life stories on LessWrong — ShakeelHashim · 2026-09-11
- Kid Coins "Princessmaxxing" After Subway Chat About Same-Sex Wedding Attire — anderssandberg · 2026-09-11