Natolambert finishes a new RLHF book built from his Olmo training experience

natolambert · x · 2026-07-21

A new RLHF book is out

Natolambert says the book Reinforcement Learning from Human Feedback is finished and is now available through the website, Amazon, and Manning.

He says it is the book he wished he had while learning to fine-tune, align, and post-train models after ChatGPT. The project was written nights and weekends since 2024, drawing on lessons from building Olmo and turning those intuitions into book form.

Related event: Nathan Lambert Completes RLHF Book After Two Years with Courses and Code(8 posts)→

Original post →

More from Research

Research channel →