Researcher Questions Benchmark Data as Newer Models Score Below Older Ones
Researcher BlackHC publicly questioned the evaluation results of a benchmark (reportedly Meta-related), noting that newer models fail to outperform older generations on the leaderboard and that other scores also appear anomalous. He demanded clarification on the benchmark's overall evaluation methodology.
2026-10-01 ~ 2026-10-01 · 2 related posts
- Researcher questions anomalous benchmark evals where new models fail to beat older ones — BlackHC · 2026-10-01
1 near-duplicate retellings: BlackHC