CritPt Physics Benchmark Tests LLM Research Reasoning, Top Scores Hit 30%
geoffwolfe · x · 2026-08-24
ReasonCore introduced CritPt (Critical Point), a benchmark designed to test frontier LLMs on real physics research capabilities. It features 71 research-scale challenges written by 50+ physicists based on unpublished work, targeting integrated reasoning rather than isolated knowledge points.
- Design: Problems are memory-resistant with machine-verifiable answers (numerical arrays, Python functions), eliminating LLM-as-judge subjectivity.
- Results: The best model scored 5.7% at launch; top models now approach 30%.
- Insight: Progress is rapid, but 70% of the map remains dark, highlighting a persistent gap in sustained, integrated reasoning.
More from Research
- Study: Internet Era Rise of System-level Creativity in Science — JMateosGarcia · 2026-08-24
- SMPLOlympics: RL Policies Trained for 10+ Humanoid Sports in Simulation — zhengyiluo · 2026-08-24
- Paper: Introduction to Simulation-Based Inference with ML — RexDouglass · 2026-08-24
- Yoav Goldberg: Stop requiring AI usage disclosures in papers — yoavgo · 2026-08-24
- AAAI 2027 Addresses Reviewer Collusion and 2-Cycles — Fragrant_Fan_6751 · 2026-08-24
- T7 Promoter Calculator predicts transcription rates accurately — anshulkundaje · 2026-08-24