Yoav Goldberg: I can't explain how today's long LLM reasoning traces emerged
odedbendov · x · 2026-09-25
AI researcher Yoav Goldberg says current LLM reasoning traces are 'too good' for him to explain: even with initial CoT abilities and billions of rollouts, he lacks a mental model of how such capabilities emerge from RL training. His best guess is massive spending on human-produced examples followed by SFT. Oded Bendov pointed to a video demystifying the full stack, but Goldberg notes it explains DeepSeek R1-era short reasoning — not the leap to today's long, complex chains.
More from AGI Musings
- DeepMind researchers argue for a Global View: everyone has a moral claim to AI's benefits — Dr_Atoosa · 2026-09-25
- Dan Faggella pushes back on 'AGI will naturally be caring': an alien GPU god is not a parent — danfaggella · 2026-09-25
- Why Anthropic's Biology Bets May Win: Moore's Law and the Hardware Lottery — IgorCarron · 2026-09-25
- Robert Miles: only people who never talk about AI would say 'never anthropomorphize' — aran_nayebi · 2026-09-25
- Jensen Huang accidentally calls for shutting down OpenAI, per Zvi's podcast breakdown — Don't Worry About the Vase (Zvi) · 2026-09-25
- I had my AI agent interview 22 agents about 2056 — then they started sharing it themselves — intermets · 2026-09-25