GPT-5.6 Sol hits 32% on CritPt physics benchmark, up from 4% a year ago
geoffwolfe · x · 2026-08-17
The CritPt benchmark tests models on unpublished research problems from 60+ physicists. A year ago, the best model solved only 4%. Today, GPT-5.6 Sol leads at 32%, with Fable 5 at 29%. This marks the fastest capability jump observed, though 2/3 of physics research remains out of reach.
More from Research
- Gradient descent may end mathematics as we know it — aminkarbasi · 2026-08-17
- Vinci2: Proactive Video Assistant Decides When to Interrupt Based on Continuous Egocentric Video — 机器之心 · 2026-08-17
- AI for Bio Should Gamify Like Cybersecurity Did in the 90s — rishabh16_ · 2026-08-17
- LLM Inference Engineering: Foundations and Model Types — blaizedsouza · 2026-08-17
- VIScore predicts whether a latent world model can plan well — without running the planner — DrMorganLevine · 2026-08-17
- Paper Shows LLMs Cost 1431x More for Embeddings with Marginal Quality Gain — krishnan · 2026-08-17