GPT-5.6 Sol hits 32% on CritPt physics benchmark, up from 4% a year ago

geoffwolfe · x · 2026-08-17

The CritPt benchmark tests models on unpublished research problems from 60+ physicists. A year ago, the best model solved only 4%. Today, GPT-5.6 Sol leads at 32%, with Fable 5 at 29%. This marks the fastest capability jump observed, though 2/3 of physics research remains out of reach.

Original post →

More from Research

Research channel →