Five years ago 50% was SOTA on math benchmarks — and AI progress keeps accelerating
minilek · x · 2026-09-10
Pushing back on claims that catastrophic AI outcomes are impossible science fiction, researcher minilek points to math benchmark history: five years ago, 50% accuracy on GSM8K-style math was state of the art for LLMs, and it took only a couple of years to saturate that benchmark entirely. His conclusion: we're already living in science fiction, and the pace is accelerating — so 'impossible' arguments don't hold.
More from AGI Musings
- David Khourshid: curiosity, pushback and rest are the human edges agents can't fake — DavidKPiano · 2026-09-10
- After Coxon's 150M-view exit warning, a case against an AI 'Patriot Act' — AIandDesign · 2026-09-10
- Defining agent delegation: what's delegated, alienability, and by whom — charles_irl · 2026-09-10
- OpenAI claims Navier-Stokes breakthrough, says another Millennium Prize problem near — basedjensen · 2026-09-10
- Agent permissions resemble human permissions: think delegation, not processes — charles_irl · 2026-09-10
- ~25% of NBER working papers contain AI text; one program hits 32%, per Pangram analysis — soumitrashukla9 · 2026-09-10