Nathan Lambert says his RLHF book is finished after two years of nights and weekends

TheZachMueller · x · 2026-07-21

Nathan Lambert says his book Reinforcement Learning from Human Feedback is done.

He describes it as the resource he wished he had while learning to fine-tune, align, and now post-train models after ChatGPT. The book has been built over nights and weekends since 2024, and he says he tried to transfer as much of the intuition from building Olmo into book form as possible.

The post frames the book as a practical foundation for people working on RLHF and post-training rather than as a polished marketing launch.

Related event: Nathan Lambert Completes RLHF Book After Two Years with Courses and Code(8 posts)→

Original post →

More from Companies & People

Companies & People channel →