"What's the Secret Sauce?" Barak Rotblat on How RL Yields Elite LLM Reasoning

BarakRotblat · x · 2026-09-25

Barak Rotblat admits in a discussion with Yoav Goldberg that current LLM reasoning traces are "too good" to explain: he has no mechanistic model of how they emerge from RL training, even with initial CoT abilities and billions of rollouts.

His best guess: enormous spending on human-produced examples followed by SFT — but he doubts that's really all. Goldberg, for his part, rejects the "emergence" framing entirely.

Related event: Yoav Goldberg Questions Whether Frontier LLM Reasoning Is Truly Emergent, Sparking Researcher Debate(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →