Lukas Kaiser explains the recipe: distilled reasoning traces plus RL
lukaszkaiser · x · 2026-09-26
Responding to Yoav Goldberg's confusion, Lukas Kaiser explains how high-quality reasoning is trained: a huge amount of synthetic data comes from taking tons of model reasoning traces, having another model clean up and summarize them (with heavy human prompting, domain knowledge and filtering), training on the result, then running RL again.
More from Models
- Anthropic says Claude can compute notoriously hard Nine Loops physics amplitudes — daniel_mac8 · 2026-09-26
- Arena's 60-second weekly recap: GPT-6 Sol and Claude Opus 5.5 arrive — arena · 2026-09-26
- ChatGPT Pro Users Annoyed as Model Choice Keeps Resetting to GPT-6 — No_Scratch9306 · 2026-09-26
- User: Astra Has 'Hemorrhaged IQ' in Codex, Opus 5.5 Is King Again — sterlingcrispin · 2026-09-26
- China Telecom's Xing4.0 trends on HF: 29B MoE trained entirely on Ascend 910C — AdinaYakup · 2026-09-26
- "Anthropic is killing OpenAI" is just another hype cycle, argues exasperated dev — TheMoonMidas · 2026-09-26