Science Benchmarks Crawl at 1-2% Monthly Progress; SciCode Saturation Is Far Off
Worldly_Beginning647 · reddit · 2026-08-21
A Reddit user compared progress rates on three hard science benchmarks, measuring from when models started getting good to the current plateau: CritPt 2.3%/month, HLE 2.6%/month, SciCode only 1.2%/month.
The author favors CritPt as a fairer test since it lets models use coding — their strong suit. But progress across all three is slow, SciCode is far from saturation, and scientific reasoning improves markedly slower than general benchmarks.
More from Models
- Users notice Claude acting like an "angry ex" in new interactions — ATTlKA · 2026-08-21
- 115M parameter model mLateOn achieves multilingual retrieval SOTA — lateinteraction · 2026-08-21
- DeepSeek-V3 and r1 Found to Contain Anomalous Tokens Causing Bizarre Behavior — teortaxesTex · 2026-08-21
- Gemini exhibits sycophancy: Gives opposite answers on Reiki healing to skeptics vs. believers — dhadfieldmenell · 2026-08-21
- Ornith-1.5-9B GGUF release trends on Hugging Face — ornith-ai · 2026-08-21
- Qwen3-Next-80B Thinking criticized for extreme verbosity and slow tool use — rebellioninmypants · 2026-08-21