Natolambert’s RL course lecture maps KL regularization to better generalization than SFT

natolambert · x · 2026-07-29

Lecture 10 of Natolambert’s course focuses on regularization in RL, especially the evolving role of the KL penalty. It also surveys papers arguing that RL can generalize better than SFT, with theory-backed explanations.

The lecture closes by connecting these ideas to a broader pattern in ML: once a control problem becomes important, the same regularization tensions tend to reappear in a different form.

Related event: RL Course Explores KL Regularization and Generalization Benefits(4 posts)→

Original post →

More from Research

Research channel →