Zero SFT needed: researcher argues superhuman reasoning emerges from pretraining plus scaled RL
brianryhuang · x · 2026-09-25
Extending the debate with Yoav Goldberg, brianryhuang argues SFT on human-curated examples only lifts the y-intercept of the scaling curve — a nice bonus but unnecessary. Superhuman reasoning should be possible with zero SFT: pretraining on 'non-agent' corpuses followed by heavily scaled RL.
The claim: frontier reasoning isn't distilled from human demonstrations; it emerges from pretraining foundations plus large-scale RL, with SFT merely a bonus.
More from Research
- Nokia Open-Sources AnyJev: Turn Any LLM into a Calibrated Decision Model, No Training — kalyan_kpl · 2026-09-25
- Why GPUs need philox, not xorshift: parallel RNG in AI training explained — abhi9u · 2026-09-25
- Independent researcher: activation steering measures the wrong geometry — 11 experiments on Qwen2.5-7B yield AkbasCore 3.2 — Nearby_Indication474 · 2026-09-25
- Navigating tenure-track in 2026: a guide to the academic job market — mboehme_ · 2026-09-25
- Grady Booch: LLMs Only Resemble the Brain at Its Most Primitive Structures — Grady_Booch · 2026-09-25
- ICLR 2027 Submission De-anonymization Incident Sparks OpenReview Statement — Striking-Warning9533 · 2026-09-25