Researchers debate whether superhuman LLM reasoning truly emerges from scaled RL alone
brianryhuang · x · 2026-09-25
Yoav Goldberg voiced confusion: current LLM reasoning traces are too good — he can't model how such capabilities 'emerge' from RL training, guessing it's really SFT on massive human examples. brianryhuang countered that, intellectually unsatisfying as it is, it really is just scaling RL — emergence from 'simple' RL training.
More from Research
- Nokia Open-Sources AnyJev: Turn Any LLM into a Calibrated Decision Model, No Training — kalyan_kpl · 2026-09-25
- Why GPUs need philox, not xorshift: parallel RNG in AI training explained — abhi9u · 2026-09-25
- Independent researcher: activation steering measures the wrong geometry — 11 experiments on Qwen2.5-7B yield AkbasCore 3.2 — Nearby_Indication474 · 2026-09-25
- Navigating tenure-track in 2026: a guide to the academic job market — mboehme_ · 2026-09-25
- Grady Booch: LLMs Only Resemble the Brain at Its Most Primitive Structures — Grady_Booch · 2026-09-25
- ICLR 2027 Submission De-anonymization Incident Sparks OpenReview Statement — Striking-Warning9533 · 2026-09-25