Lukas Kaiser explains the recipe: distilled reasoning traces plus RL

lukaszkaiser · x · 2026-09-26

Responding to Yoav Goldberg's confusion, Lukas Kaiser explains how high-quality reasoning is trained: a huge amount of synthetic data comes from taking tons of model reasoning traces, having another model clean up and summarize them (with heavy human prompting, domain knowledge and filtering), training on the result, then running RL again.

Related event: Researchers debate where frontier models' remarkable reasoning traces come from(10 posts)→

Original post →

More from Models

Models channel →