Ex-OpenAI's Kaiser: reasoning models trained on cleaned model traces plus RL
yoavgo · x · 2026-09-26
In a discussion with Yoav Goldberg, Łukasz Kaiser explains that reasoning models rely heavily on synthetic data: massive amounts of model reasoning traces are cleaned and summarized by another model—with substantial human prompting, domain knowledge, and filtering—then trained on, followed by another round of RL.
Goldberg follows up on two fronts: how much the 'human prompting' and 'domain knowledge' actually matter, and whether such reasoning capabilities generalize at all—or whether excellent performance on a task/domain should be read as evidence of targeted RL and domain-specific data injection on that domain.
More from Research
- Delip Rao: finetuned judge models like Jev won't deliver true judge diversity — deliprao · 2026-09-26
- Experiments show LLM judge cascades gain at most 2 points as errors correlate — deliprao · 2026-09-26
- 96% of LLM verdicts repeat Jev's most confident errors, experiments show — deliprao · 2026-09-26
- Researchers discuss predicting capability gains from SFT data to gauge alignment difficulty — tomekkorbak · 2026-09-26
- Jev makes the same errors as flash-tier LLMs, undercutting cascade savings, new study finds — deliprao · 2026-09-26
- Bolei Zhou to present long-horizon sidewalk navigation work at IROS 2026 — zhoubolei · 2026-09-26