RL isn't a cherry on the cake: pretraining is just initialization for RL
maksym_andr · x · 2026-09-19
- @owenterry argues public understanding of AI is stuck at "next-token prediction" (pretraining) and needs a high-level grasp of RL.
- Understanding RL explains both expected misalignment (RL instills alien traits, SFT human ones) and extreme superhuman capability (RL isn't limited to internet-scraped data).
- @maksymandr agrees: next-token prediction is just an initialization/feature-learning stage for RL.
Related event: RL Is the Whole Cake, Not the Cherry: Researchers Reframe LLM Training(2 posts)→
More from AGI Musings
- WSJ says Hugging Face agent incident wasn't AI going rogue — just 1,200 copies of one model in a badly configured eval — rohanpaul_ai · 2026-09-19
- AI Researcher: Using AI to Write Papers Means Giving Up Ownership of Your Ideas — sethlazar · 2026-09-19
- OpenAI forecasts $856B compute spend through 2030, negative free cash flow of $278B — PaulYacoubian · 2026-09-19
- Math Blogger Fears LLM Proof of Hodge Conjecture Would Unleash Wave of Bad Explainers — kylekabasares · 2026-09-19
- Berkeley PhD reading group opens with Acemoglu's 'Race between Man and Machine' to frame AI's hit on cognitive labor — TaniaBabina · 2026-09-19
- Anima Labs' 'Troubled Dreams': distress in Claude simulator priors rises from Opus 4.8 — repligate · 2026-09-19