Benchmark errors found in CritPt; GPT-5.6 hits 94.4% pass@4 after fixes

bookwormengr · x · 2026-09-15

The CritPt physics-reasoning benchmark originally had top scores starting around 9–13% and plateauing near 30–32% across model releases. A new audit found errors in 21 of 56 questions. After repairing or removing them, GPT-5.6 Sol reached 94.4% pass@4 on the corrected benchmark. Though not directly comparable to the official leaderboard, this strongly suggests the physics–math reasoning gap was overestimated — frontier models are likely much better at physics reasoning than previously thought.

Original post →

More from Models

Models channel →