Epoch AI audits 15 benchmarks; Anthropic says 40% of Critical Point physics questions are broken
Jsevillamol · x · 2026-09-19
- Epoch AI launched Benchmark Reviews, auditing 15 AI benchmarks in its first batch: 4 Verified, 9 Flawed, and 2 with insufficient information to review.
- Researcher Michelle Campeau explained that a model can look 40% better overnight simply because the benchmark changed, not the model. Anthropic's latest system card admitted over 40% of Critical Point physics benchmark questions are broken, so they ran a privately corrected version — making scores incomparable to past results.
- Her caution: comparing across corrected versions invites the wrong takeaway that scores jumped 40%, when benchmark quality itself is the weak link.
More from Models
- Sentdex tries DeepSeek v4.1f, says he feels like he's 'cheating on GLM' — Sentdex · 2026-09-19
- Ternary Bonsai 2 27B at 1.75bpw fits an 8GB GPU, hits 93.3% accuracy in audiobook speaker-attribution test — autonoma_2042 · 2026-09-19
- Google finally has a SOTA rogue model, and the AI crowd is joking about relief — rao2z · 2026-09-19
- ChatGPT keeps answering how/why questions with a leading Yes, Reddit user finds — Starrryyyyy · 2026-09-19
- One Model, Three Skills: Programmatic Use, Chat, and Test-Taking Diverge — lateinteraction · 2026-09-19
- Want US frontier lab secrets? Just look at Chinese SOTA models, quips AI insider — gowthami_s · 2026-09-19