What bridged the gap from R1-style short reasoning to today's long chains?

yoavgo · x · 2026-09-25

Yoav Goldberg adds that the known training stack roughly explains DeepSeek R1-era 'short and okay-ish' reasoning, but current models show far more impressive abilities. What bridged the gap between short reasoning then and the long, complex traces now remains unexplained — an open jab at the opacity of frontier reasoning-model training.

Related event: Yoav Goldberg Questions Whether Frontier LLM Reasoning Is Truly Emergent, Sparking Researcher Debate(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →