OpenAI's RLSlow team: how RL for reasoning earned conviction and scaled
lukaszkaiser · x · 2026-09-07
Former OpenAI researcher Lukasz Kaiser paid tribute to the RLSlow team, whose name nods to Thinking, Fast and Slow: LLMs already had "fast" thinking (immediate answers), and RL could teach them "slow" thinking — deliberate, multi-token reasoning that spends more compute. TrapitBansal recalls how early the team developed real conviction, spending long nights babysitting runs and designing careful experiments, with early discussions alongside Ilya Sutskever and later Jakub Pachocki. Evidence pushed them early to a strong view: RL for reasoning would scale. Kaiser thanks the full roster spanning the foundations of reasoning (later project "berry") onward, including Francis Song, Suchir Balaji and Dan Selsam.
Related event: Ex-OpenAI Researcher Lukasz Kaiser Bids Farewell to RLSlow Team(2 posts)→
More from Companies & People
- OpenAI researcher's sobering essay draws calls to disclose more about AI concerns — JeffLadish · 2026-09-07
- Ex-NeurIPS SAC: decline one invite and you're never asked to serve again — andrewgwils · 2026-09-07
- Caltech Mathathon sets fairness rules: all AI conversations open-sourced for math community review — FinanceYF5 · 2026-09-07
- Researcher calls undisclosed AI safety incident 'very bad,' disclosure excuse absurd — eliebakouch · 2026-09-07
- Clara Shih on when students should start using AI: once you can judge the output — clarashih · 2026-09-07
- Claude Code's Boris Cherny: don't optimize token cost, maximize returns — rohanpaul_ai · 2026-09-07