Yale physicists re-grade AI benchmarks: most are broken, frontier models near saturation

inductionheads · x · 2026-09-15

A new paper, "How Good Are Frontier Models at Physics?", had Yale physicists expertly re-grade existing benchmarks — and found many model "failures" were actually broken benchmark items.

Takeaway: most current physics benchmarks are broken and no longer measure frontier capability.

Original post →

More from Models

Models channel →