Researcher questions anomalous benchmark evals where new models fail to beat older ones

BlackHC · x · 2026-10-01

Researcher BlackHC is publicly questioning the eval numbers on a benchmark, noting that new model generations do not appear to outperform older ones at all — contrary to expectations — and that other numbers look off too. He is pressing for an explanation of what is going on with the benchmark's data overall.

Related event: Researcher Questions Benchmark Data as Newer Models Score Below Older Ones(2 posts)→

Original post →

More from Models

Models channel →