Epoch launches Benchmark Reviews, auditing 15 AI benchmarks: only 4 verified, 9 flawed
stochasticchasm · x · 2026-09-18
- Epoch AI launched Benchmark Reviews, a new initiative to audit AI benchmarks, starting with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with insufficient information for review.
- The finding that most popular benchmarks are flawed is a big deal for anyone doing agent evals.
- Community reaction: Arthur Bresnu called it "the state of agentic evals in one tweet" and hoped it would force people to take benchmark quality more seriously.
Related event: Epoch AI Audits 15 AI Benchmarks, Finds 9 Flawed(8 posts)→
More from Research
- New Paper: Contrastive Noise Alignment Cuts FID Over 50% in Few-Step Generation — burny_tech · 2026-09-18
- ICLR26 Paper Defines 'Interpretive Equivalence': Comparing Neural Network Algorithms Without Full Interpretation — burny_tech · 2026-09-18
- DeepSWE-mini: A 16-Instance Subset That Replicates the DeepSWE Leaderboard Rankings — asankhs · 2026-09-18
- Tsinghua and ByteDance Seed unveil SMELT, the first fair comparison of Looped Transformers — jiqizhixin · 2026-09-18
- Gary Marcus: Judea Pearl's causality challenges remain unsolved in the LLM paradigm — GaryMarcus · 2026-09-18
- TPA: Tensor Product Attention unifies MHA/GQA/MLA, NeurIPS 2025 Spotlight — burny_tech · 2026-09-18