"AI is only trained on human data" is outdated: RL exploration now drives frontier training
aran_nayebi · x · 2026-09-22
Jeff Ladish pushes back on the common claim that AI is only trained on human data: it was basically true when ChatGPT and GPT-4 launched, but starting in 2024, companies figured out how to use reinforcement learning to train models to explore and learn through trial and error.
He quotes Plinz's argument that many people think LLMs are limited by human training data, but pretraining is merely establishing a common-sense baseline — most compute now goes into RL, allowing models to move beyond human level by exploring themselves, just as happened with Go.
The thread is a direct rebuttal to the "models are capped by human data" critique, pointing to RL-based exploration as the mainstream frontier training approach.
Related event: Debate Erupts Over Whether LLMs Can Surpass Human-Level Intelligence(3 posts)→
More from AGI Musings
- Physicist Quits Tenure-Track Job for AI Safety: 'They Installed Escalators on All the Mountains' — matthew_d_green · 2026-09-22
- AI x-risk debate reignited: distraction conspiracy or inconvenient truth? — AaronBergman18 · 2026-09-22
- Alex Epstein calls (P)doom 'fake threat analysis' that only manufactures fear — TinfoilTricorn · 2026-09-22
- Ruxandra Teslo: AI anxiety is about losing meaning and agency, not just jobs — anshulkundaje · 2026-09-22
- From Erdős to Millennium Problems: AI math progress is compounding fast — haider1 · 2026-09-22
- More inference than training machines means writing will soon beat reading — GregoryDiamos · 2026-09-22