Nathan Lambert says his RLHF book is finished after a year of nights and weekends

alexisjross · x · 2026-07-21

Nathan Lambert says his book Reinforcement Learning from Human Feedback is finished.

He describes it as the book he wished he had while learning to fine-tune, align, and post-train models after ChatGPT, and says it was built from night-and-weekend study and documentation work since 2024.

Related event: Nathan Lambert Completes RLHF Book After Two Years with Courses and Code(8 posts)→

Original post →

More from Companies & People

Companies & People channel →