Yoav Goldberg Questions Whether Frontier LLM Reasoning Is Truly Emergent, Sparking Researcher Debate
AI researcher Yoav Goldberg raised an industry-level question on September 25: the quality of current LLM reasoning traces is so high that he cannot understand it—even with models possessing initial CoT capabilities and tens of billions of rollouts, he still cannot build a mental model explaining how these abilities "emerge" from RL training. His guess is that it relies on SFT with massive human-annotated examples, but he himself is not certain.
Confirmed
- This is Goldberg's own publicly expressed puzzlement and judgment on Twitter/X, not a secondhand report.
- Goldberg added that the existing public tech stack can roughly explain the "short but adequate" reasoning of the DeepSeek R1 generation, but what filled the gap between short reasoning and today's long, complex reasoning is something he completely cannot explain.
- brianryhuang joined the discussion, creating a clash of views (see below).
Unconfirmed
- Whether reasoning capability mainly comes from RL emergence or large-scale SFT annotation remains uncertain—Goldberg explicitly said so, and his SFT guess is only a hypothesis.
- brianryhuang's claim that "pure RL with zero SFT can achieve superhuman reasoning" is a theoretical position, not yet publicly empirically confirmed.
Why it matters
- This discussion touches the most central mechanistic question in the current reasoning-model boom: why models can reason and where the capability comes from, which directly bears on training recipes, data investment directions, and reproducibility.
- brianryhuang's view is quite representative: while the answer may be "intellectually unsatisfying," SFT on human-curated examples is merely a bonus that raises the scaling curve's intercept, not a necessity—superhuman reasoning is theoretically achievable with zero SFT. If true, it means the RL route's ceiling and cost structure may differ from mainstream beliefs.
2026-09-25 ~ 2026-09-25 · 8 related posts
Primary sources
- [source] Yoav Goldberg: LLM reasoning traces are 'too good' — unclear how they emerge from RL — yoavgo · 2026-09-25
- [source] What bridged the gap from R1-style short reasoning to today's long chains? — yoavgo · 2026-09-25
- Zero SFT needed: researcher argues superhuman reasoning emerges from pretraining plus scaled RL — brianryhuang · 2026-09-25
- "What's the Secret Sauce?" Barak Rotblat on How RL Yields Elite LLM Reasoning — BarakRotblat · 2026-09-25
- [source] Yoav Goldberg: Frontier Model Reasoning Is Not "Emergent" — yoavgo · 2026-09-25
- Yoav Goldberg: Frontier Model Reasoning Looks Like a Training Process Question, Not a Black Box — yoavgo · 2026-09-25
2 near-duplicate retellings: brianryhuang · odedbendov