Expert re-grading shows physics benchmarks are broken: most 'model errors' are benchmark or grading errors

zainhas · x · 2026-09-16

An arXiv paper with 50+ authors (including physics professors) revisits frontier LLM performance on six widely used physics benchmarks. Low reported scores — including those in the Artificial Analysis Intelligence Index (2026) — suggest models struggle with advanced physics, but expert re-grading reveals that most errors are benchmark errors or grading errors, not model errors. Several benchmarks are near saturation and can no longer reliably distinguish frontier models.

Related event: Physicists Re-evaluate Benchmarks: Frontier Models Near Physics Saturation(2 posts)→

Original post →

More from Models

Models channel →