Yoav Goldberg: I can't explain how today's long LLM reasoning traces emerged

odedbendov · x · 2026-09-25

AI researcher Yoav Goldberg says current LLM reasoning traces are 'too good' for him to explain: even with initial CoT abilities and billions of rollouts, he lacks a mental model of how such capabilities emerge from RL training. His best guess is massive spending on human-produced examples followed by SFT. Oded Bendov pointed to a video demystifying the full stack, but Goldberg notes it explains DeepSeek R1-era short reasoning — not the leap to today's long, complex chains.

Related event: Yoav Goldberg Questions Whether Frontier LLM Reasoning Is Truly Emergent, Sparking Researcher Debate(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →