DeepSeek-R1 grew reasoning with pure RL, no SFT — and that's what changed everything
McDonaghMatthew · x · 2026-09-07
- Matt McDonagh argues DeepSeek's Jan 22, 2025 R1 paper changed humanity's trajectory — not for performance or cheap training, but how it was achieved.
- DeepSeek-R1-Zero skipped initial SFT and was trained with RL directly on a base model, spontaneously developing chain-of-thought reasoning, self-verification, and reflection.
- The model showed a self-evolution process, extending its own thinking time and hitting an "aha moment" where it learned to re-evaluate its approach — evidence RL can let models discover problem-solving strategies autonomously.
- Implication: general reasoning intelligence can be organically grown without supervised data, and AI is coming for all jobs, including AI jobs.
More from AGI Musings
- Sam Altman says AI is heading to autonomous research — are firms ready? — GabrieLX5 · 2026-09-07
- Gary Marcus: AGI Talk Is a Pump to Dump IPO Stocks on Retail Investors — GaryMarcus · 2026-09-07
- Beyond the Bitter Lesson: Is There an Optimistic Sweeter Lesson? — juansequeda · 2026-09-07
- HKU Seminar Argues LLMs' Math Success Is Fragile: Inductive Engines Can't Reliably Do Deduction — YiMaTweets · 2026-09-07
- Lex Sokolin: A graveyard of early attempts is rarely proof the thesis was wrong — LexSokolin · 2026-09-07
- Researchers publish blog post on anthropomorphism in AI explanations — Dr_Atoosa · 2026-09-07