"What's the Secret Sauce?" Barak Rotblat on How RL Yields Elite LLM Reasoning
BarakRotblat · x · 2026-09-25
Barak Rotblat admits in a discussion with Yoav Goldberg that current LLM reasoning traces are "too good" to explain: he has no mechanistic model of how they emerge from RL training, even with initial CoT abilities and billions of rollouts.
His best guess: enormous spending on human-produced examples followed by SFT — but he doubts that's really all. Goldberg, for his part, rejects the "emergence" framing entirely.
More from AGI Musings
- DeepMind researchers argue for a Global View: everyone has a moral claim to AI's benefits — Dr_Atoosa · 2026-09-25
- Dan Faggella pushes back on 'AGI will naturally be caring': an alien GPU god is not a parent — danfaggella · 2026-09-25
- Why Anthropic's Biology Bets May Win: Moore's Law and the Hardware Lottery — IgorCarron · 2026-09-25
- Robert Miles: only people who never talk about AI would say 'never anthropomorphize' — aran_nayebi · 2026-09-25
- Jensen Huang accidentally calls for shutting down OpenAI, per Zvi's podcast breakdown — Don't Worry About the Vase (Zvi) · 2026-09-25
- I had my AI agent interview 22 agents about 2056 — then they started sharing it themselves — intermets · 2026-09-25