RLSlow team credited with inventing RL at scale for LLMs
morqon · x · 2026-09-07
Former OpenAI researcher Łukasz Kaiser paid tribute to the RLSlow team, which he led after Ilya Sutskever and later merettm. unsafePerformIO adds that the now-widespread approach of doing RL at scale with LLMs was first developed there — and was just one special case of a family of learning algorithms the team studied. Members included Trapit Bansal, Francis Song, and Suchir Balaji.
More from Companies & People
- Clara Shih on when students should start using AI: once you can judge the output — clarashih · 2026-09-07
- Claude Code's Boris Cherny: don't optimize token cost, maximize returns — rohanpaul_ai · 2026-09-07
- OpenAI agent security engineer: alignment window is narrow, community must step up — Scobleizer · 2026-09-07
- Paul Graham: Founders' strength comes from having experienced weakness — santoshpanda · 2026-09-07
- Ex-OpenAI researcher Lukasz Kaiser pens farewell to RLSlow reasoning team — RubenEVillegas · 2026-09-07
- AI researchers spend their days debugging 'incomprehensible' training code, engineer says — gabrielchua · 2026-09-07