Epoch AI audits 15 AI benchmarks: 4 Verified, 9 Flawed; PostTrainBench passes and preps v1.2
maksym_andr · x · 2026-09-18
Epoch AI launched Benchmark Reviews, a new initiative auditing AI benchmarks, starting with 15: 4 Verified, 9 Flawed, and 2 with insufficient information. PostTrainBench v1.1 was designated Verified, and its team published a point-by-point response:
- On narrow overfitting to a single eval: this is intended — each agent gets one H100 for ten hours of post-training, so teaching broad transferable skills is unrealistic; the question is whether the agent improves the model on one target benchmark
- On middling averages and high variance: flagged runs are manually verified to keep the judge calibrated
- The team agrees with some findings and will release v1.2 addressing the flagged issues plus new ones found internally
More from Models
- Sentdex tries DeepSeek v4.1f, says he feels like he's 'cheating on GLM' — Sentdex · 2026-09-19
- Ternary Bonsai 2 27B at 1.75bpw fits an 8GB GPU, hits 93.3% accuracy in audiobook speaker-attribution test — autonoma_2042 · 2026-09-19
- Google finally has a SOTA rogue model, and the AI crowd is joking about relief — rao2z · 2026-09-19
- ChatGPT keeps answering how/why questions with a leading Yes, Reddit user finds — Starrryyyyy · 2026-09-19
- One Model, Three Skills: Programmatic Use, Chat, and Test-Taking Diverge — lateinteraction · 2026-09-19
- Want US frontier lab secrets? Just look at Chinese SOTA models, quips AI insider — gowthami_s · 2026-09-19