Stop Calling AI 'Trained on Human Data': RLVR Is Where the Real Gains Come From

burny_tech · x · 2026-09-23

A pushback against the outdated framing that AI is merely 'trained on human data': pretraining uses human data, but modern models are increasingly shaped by RLVR (RL with verifiable rewards), which is decidedly not human data — solutions to Navier-Stokes and other Millennium Problems aren't in any internet corpus.

The author argues that modeling AI as 'a thing trained on what has already been written' makes you miss its most important capabilities. A quoted reply counters that since pretraining data contains nothing smarter than humans, progress may stall near human-level intelligence — a core disagreement about current AI training paradigms.

Original post →

More from AGI Musings

AGI Musings channel →