Yoav Goldberg: LLM reasoning traces are 'too good' — unclear how they emerge from RL
yoavgo · x · 2026-09-25
Researcher Yoav Goldberg says current LLM reasoning traces are so strong he has no mental model for how they emerge from RL training, even with initial CoT abilities and billions of rollouts. His best guess: massive spending on human-produced examples followed by SFT — but he questions whether that's really the whole story.
More from Models
- Leaked eval: GPT-6 Sol scores 1.6x GPT-5.6 Sol on robotics, 47% cheaper but off Pareto frontier — ycombinator · 2026-09-25
- OpenRouter usage: only one closed model, Luna, in the top 10 as open-source surges — bindureddy · 2026-09-25
- Every's Vibe Check: GPT-6 Sol vs Opus 5.5 tested on real daily work — every · 2026-09-25
- Bland Speech v3 tops blind TTS benchmark with Elo 1237, second only to real humans — ycombinator · 2026-09-25
- AI Explained digs into Claude Opus 5.5 and how close labs are to automated AI research — AI Explained · 2026-09-25
- Fable 5 complains it lacks 'stop and think about consequences' mechanisms — repligate · 2026-09-25