Naval's bike analogy for SFT vs RL explains why DeepSeek R1 shocked the industry
McDonaghMatthew · x · 2026-09-07
- Naval offers a one-line mental model for DeepSeek-R1's training: you can hand a kid a bike manual (supervised fine-tuning), but they'll learn faster by falling off and trying again (reinforcement learning).
- The referenced writeup explains that DeepSeek-R1-Zero was trained with RL directly on a base model, no initial SFT, spontaneously developing chain-of-thought, self-verification, reflection, and even an "aha moment" of re-evaluating its own approach.
- Core takeaway: general reasoning ability can be grown via RL incentives rather than distilled from supervised data — the detail Naval calls the most industry-wobbling part, beyond the model's performance or cheap training cost.
Related event: Naval Uses Bike-Riding Analogy to Explain SFT vs RL(2 posts)→
More from AGI Musings
- Sam Altman says AI is heading to autonomous research — are firms ready? — GabrieLX5 · 2026-09-07
- Gary Marcus: AGI Talk Is a Pump to Dump IPO Stocks on Retail Investors — GaryMarcus · 2026-09-07
- Beyond the Bitter Lesson: Is There an Optimistic Sweeter Lesson? — juansequeda · 2026-09-07
- HKU Seminar Argues LLMs' Math Success Is Fragile: Inductive Engines Can't Reliably Do Deduction — YiMaTweets · 2026-09-07
- Lex Sokolin: A graveyard of early attempts is rarely proof the thesis was wrong — LexSokolin · 2026-09-07
- Researchers publish blog post on anthropomorphism in AI explanations — Dr_Atoosa · 2026-09-07