Two years from o1-preview to superhuman math: RL scaling now cracks open research problems
jam3scampbell · x · 2026-09-09
jam3scampbell notes it took just two years to go from o1-preview's AIME test-time scaling (Sept 2024) to superhuman performance on an evaluation set of open math problems. He argues GPT-6 Astra's 'internal model' already shows a step-change just days after release, and that honest extrapolation points to fully automated research in the not-too-distant future — urging people to internalize this rate of progress and act with urgency.
More from AGI Musings
- Frontier labs quietly building recursive self-improvement, thread claims — 0xsachi · 2026-09-09
- Prediction: Big AI will start buying Big Pharma — not for the drugs, for the data — MannyKayy · 2026-09-09
- A magic lamp thought experiment on when AI solving math helps and hurts — sandersted · 2026-09-09
- Blogger: Million-Instance AI Swarms Could Cripple National Economies, and We've Seen Nothing Yet — scaling01 · 2026-09-09
- "AI Techbros Only Produce Low-Variance Copies of Researchers," Sparks Debate — MoonL88537 · 2026-09-09
- GPT-6 Astra rumored to hit ~11.6-hour p80 task horizon, tracking AI-2027 curve — haider1 · 2026-09-09