"RL is the whole cake, not the cherry": follow-up on pretraining vs reinforcement learning
maksym_andr · x · 2026-09-19
- Follow-up to @maksymandr's earlier point: against Yann LeCun's old analogy that RL is the "cherry on the cake," he argues RL is actually the whole cake.
- Next-token prediction is merely an initialization/feature-learning stage for RL, a framing he thinks should be communicated to the public.
Related event: RL Is the Whole Cake, Not the Cherry: Researchers Reframe LLM Training(2 posts)→
More from AGI Musings
- AI Researcher: Using AI to Write Papers Means Giving Up Ownership of Your Ideas — sethlazar · 2026-09-19
- OpenAI forecasts $856B compute spend through 2030, negative free cash flow of $278B — PaulYacoubian · 2026-09-19
- Math Blogger Fears LLM Proof of Hodge Conjecture Would Unleash Wave of Bad Explainers — kylekabasares · 2026-09-19
- Berkeley PhD reading group opens with Acemoglu's 'Race between Man and Machine' to frame AI's hit on cognitive labor — TaniaBabina · 2026-09-19
- Anima Labs' 'Troubled Dreams': distress in Claude simulator priors rises from Opus 4.8 — repligate · 2026-09-19
- Fine-tuned models produce far darker completions than DeepSeek V3 base, study shows — repligate · 2026-09-19