OpenAI's RLSlow team: how RL for reasoning earned conviction and scaled

lukaszkaiser · x · 2026-09-07

Former OpenAI researcher Lukasz Kaiser paid tribute to the RLSlow team, whose name nods to Thinking, Fast and Slow: LLMs already had "fast" thinking (immediate answers), and RL could teach them "slow" thinking — deliberate, multi-token reasoning that spends more compute. TrapitBansal recalls how early the team developed real conviction, spending long nights babysitting runs and designing careful experiments, with early discussions alongside Ilya Sutskever and later Jakub Pachocki. Evidence pushed them early to a strong view: RL for reasoning would scale. Kaiser thanks the full roster spanning the foundations of reasoning (later project "berry") onward, including Francis Song, Suchir Balaji and Dan Selsam.

Related event: Ex-OpenAI Researcher Lukasz Kaiser Bids Farewell to RLSlow Team(2 posts)→

Original post →

More from Companies & People

Companies & People channel →