Researchers debate whether superhuman LLM reasoning truly emerges from scaled RL alone

brianryhuang · x · 2026-09-25

Yoav Goldberg voiced confusion: current LLM reasoning traces are too good — he can't model how such capabilities 'emerge' from RL training, guessing it's really SFT on massive human examples. brianryhuang countered that, intellectually unsatisfying as it is, it really is just scaling RL — emergence from 'simple' RL training.

Related event: Yoav Goldberg Questions Whether Frontier LLM Reasoning Is Truly Emergent, Sparking Researcher Debate(8 posts)→

Original post →

More from Research

Research channel →