Natolambert finishes a new RLHF book built from his Olmo training experience
natolambert · x · 2026-07-21
A new RLHF book is out
Natolambert says the book Reinforcement Learning from Human Feedback is finished and is now available through the website, Amazon, and Manning.
He says it is the book he wished he had while learning to fine-tune, align, and post-train models after ChatGPT. The project was written nights and weekends since 2024, drawing on lessons from building Olmo and turning those intuitions into book form.
Related event: Nat Lambert Completes RLHF Book with 10+ Hour Course(11 posts)→
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11