OpenAI's Łukasz Kaiser on RL generalization: agent reasoning may generalize more than expected

lukaszkaiser · x · 2026-09-26

In a discussion with Yoav Goldberg, OpenAI researcher Łukasz Kaiser addressed whether strong model reasoning reflects targeted RL training or genuine generalization. His framing: what matters is finding domains with "enough good diverse tasks and a grader good enough" — arXiv lemmas being a good example.

Based on hands-on agent usage, he believes the models are "generalizing a fair bit," though he jokes he might just be a boring user. A substantive exchange on RL specialization vs. general reasoning.

Related event: Researchers Debate How LLMs Acquire Their Remarkable Long Reasoning Chains(12 posts)→

Original post →

More from Models

Models channel →