Most compute now goes to RL, letting models surpass human data limits

MarvinTBaumann · x · 2026-09-22

A debate over LLM ceilings: @Drachs1978 argues models will converge on human intelligence since pretraining data contains nothing smarter than humans. Plinz counters that pretraining only establishes a common-sense baseline—most compute now goes to reinforcement learning, letting models move beyond human level through self-exploration, as happened with Go. The dispute hinges on whether RL self-exploration can break past the training-data ceiling.

Related event: Debate: will LLMs plateau at human-level intelligence(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →