Stop Calling AI 'Trained on Human Data': RLVR Is Where the Real Gains Come From
burny_tech · x · 2026-09-23
A pushback against the outdated framing that AI is merely 'trained on human data': pretraining uses human data, but modern models are increasingly shaped by RLVR (RL with verifiable rewards), which is decidedly not human data — solutions to Navier-Stokes and other Millennium Problems aren't in any internet corpus.
The author argues that modeling AI as 'a thing trained on what has already been written' makes you miss its most important capabilities. A quoted reply counters that since pretraining data contains nothing smarter than humans, progress may stall near human-level intelligence — a core disagreement about current AI training paradigms.
More from AGI Musings
- Why OpenAI's vertical push into law and finance is 'dead on arrival' — nicolechirps · 2026-09-23
- '50,000 AI agents can work around any materials science bottleneck' — teortaxesTex · 2026-09-23
- Mark Cuban: healthcare benefit costs will fire more people than AI — alvelda · 2026-09-23
- 150M people alive were born before the atom bomb — why past tech says little about AI risk — AndyMasley · 2026-09-23
- One more AGI bar: replacing median white-collar work AND months-long reliability — ChrisGPT · 2026-09-23
- Dustin Moskovitz: the single largest funder of the EA ecosystem, with billions to AI safety — SydSteyerhart · 2026-09-23