Physicists Re-evaluate Benchmarks: Frontier Models Near Physics Saturation
A Yale-led paper with 50+ expert authors re-evaluated physics benchmarks and found most reported model failures were caused by flawed questions or grading, with frontier models now near saturation.
2026-09-15 ~ 2026-09-16 · 3 related posts
- Yale physicists re-grade AI benchmarks: most are broken, frontier models near saturation — inductionheads · 2026-09-15
- Most benchmark 'model errors' are actually benchmark or grading errors, analysis finds — zainhas · 2026-09-16
- Expert re-grading shows physics benchmarks are broken: most 'model errors' are benchmark or grading errors — zainhas · 2026-09-16