Open benchmarks are broken: models regex out offloaded answers, researcher argues
a1zhang · x · 2026-09-23
a1zhang argues open benchmarks have become frustrating because it's hard to tell what models actually know. On 'new' long-context benchmarks, some RLMs simply regex for offloaded information and grab the solution.
- Interpreting the HarnessTax paper: on open benchmarks, performance is mostly a function of model capability, not the harness — yet in practice different harnesses show very different behavior, exposing benchmark contamination
- Newer benchmarks constantly get swallowed, making capability gauging 'largely uninteresting'
- He rejects '2-day task' benchmarks as the answer and calls for better ways to host and manage open benchmarks that stay valid for at least a year
More from Models
- Testing the Jeb chatbot: inconsistently biased, not neutral — calibrate it like any classifier — PawarBI · 2026-09-24
- Pokemon benchmark Paradigm 3: Astra generalizes to scrambled maps and fan-made games while rivals memorize — gleech · 2026-09-24
- AI Completes Fan-Made Pokemon Brown in 10K Steps: Real Generalization or Whack-a-Mole? — gleech · 2026-09-24
- Next-gen model names surface: Opus 5.5, Fable 5.1, GPT-6 Astra — labs said to be ~2 months ahead internally — haider1 · 2026-09-24
- AI detector debate: economist argues Pangram is the only reliable tool, cites 0 FPR finding — paulnovosad · 2026-09-24
- MiniMax H3 on Spectrum Runs Each Step Twice, Killing Performance — Glittering-Cold-2981 · 2026-09-24