OpenAI's Łukasz Kaiser on RL generalization: agent reasoning may generalize more than expected
lukaszkaiser · x · 2026-09-26
In a discussion with Yoav Goldberg, OpenAI researcher Łukasz Kaiser addressed whether strong model reasoning reflects targeted RL training or genuine generalization. His framing: what matters is finding domains with "enough good diverse tasks and a grader good enough" — arXiv lemmas being a good example.
Based on hands-on agent usage, he believes the models are "generalizing a fair bit," though he jokes he might just be a boring user. A substantive exchange on RL specialization vs. general reasoning.
Related event: Researchers Debate How LLMs Acquire Their Remarkable Long Reasoning Chains(12 posts)→
More from Models
- Observation: Astra uses filler tokens far more effectively than other models — scaling01 · 2026-09-26
- How Long Until Local ~30B A3B Models Match GLM 5.3 Flash Quality? — Aggravating-Push-207 · 2026-09-26
- Ethan Mollick: 'Keep prompts short' is bad advice, and minimizing token cost confuses inputs with outputs — emollick · 2026-09-26
- Matthew Berman Reviews DeepSeek and Calls It 'Crazy' — Matthew Berman · 2026-09-26
- GPT-6 Luna uses fewer reasoning tokens than 5.6 on ARC-AGI-2, hard tasks stymie both — mhmazur · 2026-09-26
- Heavy user: fast, cheap Claude Opus 5.5 now takes all my serious work — brandon_galang · 2026-09-26