Science Benchmarks Crawl at 1-2% Monthly Progress; SciCode Saturation Is Far Off

Worldly_Beginning647 · reddit · 2026-08-21

A Reddit user compared progress rates on three hard science benchmarks, measuring from when models started getting good to the current plateau: CritPt 2.3%/month, HLE 2.6%/month, SciCode only 1.2%/month.

The author favors CritPt as a fairer test since it lets models use coding — their strong suit. But progress across all three is slow, SciCode is far from saturation, and scientific reasoning improves markedly slower than general benchmarks.

Original post →

More from Models

Models channel →