Natolambert finishes a new RLHF book built from his Olmo training experience
natolambert · x · 2026-07-21
A new RLHF book is out
Natolambert says the book Reinforcement Learning from Human Feedback is finished and is now available through the website, Amazon, and Manning.
He says it is the book he wished he had while learning to fine-tune, align, and post-train models after ChatGPT. The project was written nights and weekends since 2024, drawing on lessons from building Olmo and turning those intuitions into book form.
Related event: Nathan Lambert Completes RLHF Book After Two Years with Courses and Code(8 posts)→
More from Research
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- ICML Tutorial: Is Optimization Theory Relevant in 2026? — srush_nlp · 2026-07-22
- Reddit points to OpenAI’s ChatGPT Ads page — EcstaticAsparagus509 · 2026-07-22
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- DeepSWE: A New Benchmark for Evaluating AI Coding Agents on Real GitHub Issues — pmz · 2026-07-22