Researcher Questions Benchmark Data as Newer Models Score Below Older Ones

Researcher BlackHC publicly questioned the evaluation results of a benchmark (reportedly Meta-related), noting that newer models fail to outperform older generations on the leaderboard and that other scores also appear anomalous. He demanded clarification on the benchmark's overall evaluation methodology.

2026-10-01 ~ 2026-10-01 · 2 related posts

1 near-duplicate retellings: BlackHC