Nat Lambert finishes a RLHF book built from nights and weekends since 2024

danielhanchen · x · 2026-07-22

Nat Lambert says his book Reinforcement Learning from Human Feedback is finished, describing it as the guide he wished he had when learning to fine-tune, align, and post-train models after ChatGPT. He says the book was assembled through nights-and-weekends study since 2024, drawing on lessons from building Olmo.

Related event: Nat Lambert Completes RLHF Book with 10+ Hour Course(11 posts)→

Original post →

More from Companies & People

Companies & People channel →