Expert re-grading shows physics benchmarks are broken: most 'model errors' are benchmark or grading errors
zainhas · x · 2026-09-16
An arXiv paper with 50+ authors (including physics professors) revisits frontier LLM performance on six widely used physics benchmarks. Low reported scores — including those in the Artificial Analysis Intelligence Index (2026) — suggest models struggle with advanced physics, but expert re-grading reveals that most errors are benchmark errors or grading errors, not model errors. Several benchmarks are near saturation and can no longer reliably distinguish frontier models.
Related event: Physicists Re-evaluate Benchmarks: Frontier Models Near Physics Saturation(2 posts)→
More from Models
- OpenRouter spend flips to OpenAI over Anthropic for first time in 2.5 years — firstadopter · 2026-09-16
- KD in mid-training favors reasoning over factual recall, AI2/UW paper finds; Switch Distillation proposed — LukeZettlemoyer · 2026-09-16
- DoorDash, Siemens, Airbnb shift to cheap Chinese open-weight models — carlbfrey · 2026-09-16
- StepFun launches StepAudio 3: five audio models topping realtime voice leaderboards — StepFun_ai · 2026-09-16
- Abacus.AI says Smaug Flash fixes open-source models' tool-call hangs in production — bindureddy · 2026-09-16
- Voice mode breaking up, Codex erroring: reliability still far from AGI-ready — koltregaskes · 2026-09-16