Researcher presses for answers on benchmark evals where new models don't beat old ones
BlackHC · x · 2026-10-01
Researcher BlackHC follows up questioning a benchmark's eval numbers: his expectation is that new model generations should clearly outperform older ones, yet the numbers show otherwise, with other scores also looking anomalous. He is directly pressing the benchmark maintainers on what is going on with the data overall.
Related event: Researcher Questions Benchmark Data as Newer Models Score Below Older Ones(2 posts)→
More from Models
- Musician asks Suno and SoundCloud why its AI matched his private unreleased track — TheMoonMidas · 2026-10-01
- Early Gemini 4 Argon buzz cools as users flag poor token efficiency — adonis_singh · 2026-10-01
- Polymarket puts Google at 40% to lead AI race as Gemini 4 Argon reportedly gets 1M-token output — Polymarket · 2026-10-01
- Opus 5.5 shows a recurring 'stamp' phrase pattern, dubbed the new 'delve' — dylfreed · 2026-10-01
- Pedro Domingos: token-based pricing is doomed because tokens will soon be obsolete — pmddomingos · 2026-10-01
- ChatGPT suffers major outage, many users report service down — Polymarket · 2026-10-01