Ex-OpenAI's Kaiser: reasoning models trained on cleaned model traces plus RL

yoavgo · x · 2026-09-26

In a discussion with Yoav Goldberg, Łukasz Kaiser explains that reasoning models rely heavily on synthetic data: massive amounts of model reasoning traces are cleaned and summarized by another model—with substantial human prompting, domain knowledge, and filtering—then trained on, followed by another round of RL.

Goldberg follows up on two fronts: how much the 'human prompting' and 'domain knowledge' actually matter, and whether such reasoning capabilities generalize at all—or whether excellent performance on a task/domain should be read as evidence of targeted RL and domain-specific data injection on that domain.

Related event: Researchers debate where frontier LLM reasoning comes from; Kaiser points to massive synthetic data(11 posts)→

Original post →

More from Research

Research channel →