Epoch AI audits AI benchmarks: 9 of 15 reviewed benchmarks found flawed
joecole · x · 2026-09-18
Epoch AI launched Benchmark Reviews, a new initiative to audit AI benchmarks, starting with 15 benchmarks: only 4 are Verified, 9 are Flawed, and 2 lack sufficient information for review. The finding raises serious questions about the reliability of widely cited benchmark scores.
Related event: Epoch AI's Benchmark Reviews Flags 9 of 15 AI Benchmarks as Flawed(10 posts)→
More from Research
- Mathematicians: OpenAI's Navier-Stokes Blow-Up Construction Fails in the Zero Forcing Case — _onionesque · 2026-09-18
- Reuters: Anthropic Quietly Sets Up Bay Area Wet Lab to Let Claude Run Biology Experiments — Polymarket · 2026-09-18
- ToolUniverse hits 1M downloads, v1.5 adds structure prediction and local ML tools — marinkazitnik · 2026-09-18
- New paper applies multiscale emergence methods to EEG across conscious states — anilkseth · 2026-09-18
- Generate Your Own Erdős-Number-Style Calculator From arXiv Coauthorship Data — mircomusolesi · 2026-09-18
- ASE Submissions Doubled to 1,181 in a Year: CMU's Rohan Padhye on Why Peer Review Is Broken — moarbugs · 2026-09-18