A new RLHF book is finished after nights and weekends since 2024

HamelHusain · x · 2026-07-21

@natolambert says his book Reinforcement Learning from Human Feedback is finished. He describes it as the book he wished he had while learning to fine-tune, align, and post-train models after ChatGPT, and says it distills fundamentals he learned while building Olmo on nights and weekends since 2024.

Related event: Nathan Lambert Completes RLHF Book After Two Years with Courses and Code(8 posts)→

Original post →

More from Research

Research channel →