Benchmark errors found in CritPt; GPT-5.6 hits 94.4% pass@4 after fixes
bookwormengr · x · 2026-09-15
The CritPt physics-reasoning benchmark originally had top scores starting around 9–13% and plateauing near 30–32% across model releases. A new audit found errors in 21 of 56 questions. After repairing or removing them, GPT-5.6 Sol reached 94.4% pass@4 on the corrected benchmark. Though not directly comparable to the official leaderboard, this strongly suggests the physics–math reasoning gap was overestimated — frontier models are likely much better at physics reasoning than previously thought.
More from Models
- Google details multilingual AI push: on-device TranslateGemma and 7,000+ language data maps — ymatias · 2026-09-16
- Google's language AI spans 300+ languages reaching 7B people, unveils new research — ymatias · 2026-09-16
- Engineer explains RLHF: humans rating model outputs is standard practice at every lab — JFPuget · 2026-09-16
- ChatGPT + Gemini hold 81% of AI research share; Claude has lowest NPS, says G2 — sanderssays · 2026-09-16
- A foundation model that outputs probabilities, not words: 25x faster, ~600x cheaper — danshipper · 2026-09-16
- InternLM's Atria-Dawn-Preview, a GLM MoE-DSA model, trends on Hugging Face — internlm · 2026-09-16