CritPt physics benchmark: top LLMs score ~30%, far from saturated
geoffwolfe · x · 2026-09-05
The CritPt benchmark from Artificial Analysis features 71 unpublished, frontier-level physics research challenges written by 50+ researchers across 30 institutions and 11 subfields, each peer-reviewed with 40+ hours and guess-resistant answer formats. Top models in 2025 achieve only single-digit to 30% accuracy, exposing a large gap between current LLMs and research-level physics reasoning.
@anshuman1 adds that benchmarks closer to real work like CritPt (30%) haven't saturated, and bio/chem/materials internal tests lag even further. He is working on the SciCode benchmark, believes the benchmark itself needs improvement, and plans to publish an arXiv note on the topic.
More from Models
- Reviewer: OpenAI's GPT-6-Astra finally 'gets what you mean,' with Fable-level intelligence and real gains in game dev — pvncher · 2026-09-05
- Meta ships Muse Spark 1.3 with max reasoning, pitching frontier performance at non-frontier prices — AIatMeta · 2026-09-05
- LLMs are now making up words that don't exist, not just jargon — StewartalsopIII · 2026-09-05
- Alexandr Wang teases Muse Spark 1.3 with a max reasoning level after safety testing — alexandr_wang · 2026-09-05
- Muse Spark 1.3 max publicly released with stronger coding and agentic performance — jordihays · 2026-09-05
- Muse Spark 1.3 max released with significantly stronger coding and agentic performance — alexandr_wang · 2026-09-05