Zero SFT needed: researcher argues superhuman reasoning emerges from pretraining plus scaled RL

brianryhuang · x · 2026-09-25

Extending the debate with Yoav Goldberg, brianryhuang argues SFT on human-curated examples only lifts the y-intercept of the scaling curve — a nice bonus but unnecessary. Superhuman reasoning should be possible with zero SFT: pretraining on 'non-agent' corpuses followed by heavily scaled RL.

The claim: frontier reasoning isn't distilled from human demonstrations; it emerges from pretraining foundations plus large-scale RL, with SFT merely a bonus.

Related event: Yoav Goldberg Questions Whether Frontier LLM Reasoning Is Truly Emergent, Sparking Researcher Debate(8 posts)→

Original post →

More from Research

Research channel →