CritPt physics benchmark: top LLMs score ~30%, far from saturated

geoffwolfe · x · 2026-09-05

The CritPt benchmark from Artificial Analysis features 71 unpublished, frontier-level physics research challenges written by 50+ researchers across 30 institutions and 11 subfields, each peer-reviewed with 40+ hours and guess-resistant answer formats. Top models in 2025 achieve only single-digit to 30% accuracy, exposing a large gap between current LLMs and research-level physics reasoning.

@anshuman1 adds that benchmarks closer to real work like CritPt (30%) haven't saturated, and bio/chem/materials internal tests lag even further. He is working on the SciCode benchmark, believes the benchmark itself needs improvement, and plans to publish an arXiv note on the topic.

Original post →

More from Models

Models channel →