Yoav Goldberg: LLM reasoning traces are 'too good' — unclear how they emerge from RL

yoavgo · x · 2026-09-25

Researcher Yoav Goldberg says current LLM reasoning traces are so strong he has no mental model for how they emerge from RL training, even with initial CoT abilities and billions of rollouts. His best guess: massive spending on human-produced examples followed by SFT — but he questions whether that's really the whole story.

Original post →

More from Models

Models channel →