Researcher presses for answers on benchmark evals where new models don't beat old ones

BlackHC · x · 2026-10-01

Researcher BlackHC follows up questioning a benchmark's eval numbers: his expectation is that new model generations should clearly outperform older ones, yet the numbers show otherwise, with other scores also looking anomalous. He is directly pressing the benchmark maintainers on what is going on with the data overall.

Related event: Researcher Questions Benchmark Data as Newer Models Score Below Older Ones(2 posts)→

Original post →

More from Models

Models channel →