Naval Amplifies DeepSeek Explainer: SFT Is a Bike Manual, RL Is Learning to Ride
McDonaghMatthew · x · 2026-09-06
Naval quoted Matthew McDonagh's ELI5 explainer on DeepSeek: "Imagine teaching a child to ride a bike. You could give them a detailed manual (Supervised Fine Tuning), but they'll likely learn better by trying it themselves (Reinforcement Learning), falling, getting up, and gradually improving." McDonagh celebrated being quoted, urging everyone to keep writing words and analogies that make hard technical concepts click.
More from Models
- GPT-6 Astra Draws a Unicorn for $0.11, 15x Cheaper Than GPT-5.4 Pro — zakelfassi · 2026-09-06
- GPT-6 Astra hits 97.2% on short-doc extraction SOTA but only 31.7% on long docs — llama_index · 2026-09-06
- HF Engineer: Circuit Board Design Is My Personal Turing Test for AI — and Models Are Catching Up — hugs · 2026-09-06
- Google Astra timeouts still burn Plus-plan credits, user reports — sharkbaitlol · 2026-09-06
- Astra Minecraft test shows high reasoning one-shots a whole house via command block — hankberger · 2026-09-06
- Astra voice mode a huge step up, works well as orchestrator — jdjohnson · 2026-09-06