Yale physicists re-grade AI benchmarks: most are broken, frontier models near saturation
inductionheads · x · 2026-09-15
A new paper, "How Good Are Frontier Models at Physics?", had Yale physicists expertly re-grade existing benchmarks — and found many model "failures" were actually broken benchmark items.
- By current benchmarks, frontier models look weak at physics (e.g. GPT5.6-sol at 47.3% on Humanity's Last Exam physics), but this contradicts physicists' hands-on experience
- Re-grading the "failed" answers with Yale physicists showed the problems were usually the benchmark, not the model
- Once corrected, models almost saturate every benchmark tested, including HLE physics and CritPt
Takeaway: most current physics benchmarks are broken and no longer measure frontier capability.
More from Models
- TabPFN-3.5 launches with SOTA on complex tabular data, up to 6x faster inference — FrankRHutter · 2026-09-15
- A less boring way to browse cloud model benchmarks — the3dwin · 2026-09-15
- Flam's 26B MoE Falcon model returns first token in 30ms, specialized for Indic languages — testingcatalog · 2026-09-15
- OPEN-1B: the world's first fully auditable 1.6B-parameter transformer training run — benfielding · 2026-09-15
- xAI Reportedly Giving Out $50,000 in Grok API Credits at Random — Kyrannio · 2026-09-15
- Same feature, 3 models: $4 Claude vs $0.50 Tencent Hy3 — the cheapest won — tejasghutukade · 2026-09-15