"Don't trust benchmarks" is no excuse when 99% of them show zero improvement or regression
Angaisb_ · x · 2026-09-23
Angaisb criticizes the common "don't trust benchmarks, try it yourself" deflection: questioning benchmarks is fine when results are mixed, but when 99% of them show zero improvement or even regression, the excuse stops being serious.
More from Models
- Matt Shumer declares 'Anthropic has won,' calls new model incredible — mattshumer_ · 2026-09-24
- Testing the Jeb chatbot: inconsistently biased, not neutral — calibrate it like any classifier — PawarBI · 2026-09-24
- Pokemon benchmark Paradigm 3: Astra generalizes to scrambled maps and fan-made games while rivals memorize — gleech · 2026-09-24
- AI Completes Fan-Made Pokemon Brown in 10K Steps: Real Generalization or Whack-a-Mole? — gleech · 2026-09-24
- Next-gen model names surface: Opus 5.5, Fable 5.1, GPT-6 Astra — labs said to be ~2 months ahead internally — haider1 · 2026-09-24
- AI detector debate: economist argues Pangram is the only reliable tool, cites 0 FPR finding — paulnovosad · 2026-09-24