Researcher questions anomalous benchmark evals where new models fail to beat older ones
BlackHC · x · 2026-10-01
Researcher BlackHC is publicly questioning the eval numbers on a benchmark, noting that new model generations do not appear to outperform older ones at all — contrary to expectations — and that other numbers look off too. He is pressing for an explanation of what is going on with the benchmark's data overall.
Related event: Researcher Questions Benchmark Data as Newer Models Score Below Older Ones(2 posts)→
More from Models
- Musician asks Suno and SoundCloud why its AI matched his private unreleased track — TheMoonMidas · 2026-10-01
- Early Gemini 4 Argon buzz cools as users flag poor token efficiency — adonis_singh · 2026-10-01
- Polymarket puts Google at 40% to lead AI race as Gemini 4 Argon reportedly gets 1M-token output — Polymarket · 2026-10-01
- Opus 5.5 shows a recurring 'stamp' phrase pattern, dubbed the new 'delve' — dylfreed · 2026-10-01
- Pedro Domingos: token-based pricing is doomed because tokens will soon be obsolete — pmddomingos · 2026-10-01
- ChatGPT suffers major outage, many users report service down — Polymarket · 2026-10-01