Nat Lambert says his RLHF book is finished after nights-and-weekends writing
deliprao · x · 2026-07-21
A new RLHF book from Nat Lambert is finished and headed for launch
The post says Nat Lambert has completed Reinforcement Learning from Human Feedback, a book he says he wished existed when he was learning to fine-tune, align, and post-train models after ChatGPT.
- He built it through nights and weekends starting in 2024.
- The book is meant to capture the intuitions behind fine-tuning, alignment, and post-training.
- He says it distills lessons from building Olmo into book form.
Related event: Nathan Lambert Completes RLHF Book After Two Years with Courses and Code(8 posts)→
More from Companies & People
- Abridge’s Shiv Rao says healthcare AI wins by staying human and staying focused — elizabeth · 2026-07-22
- Abridge CEO: Domain Expertise Will Matter More in the AI Era — elizabeth · 2026-07-22
- Skyfall AI plans to buy a SaaS business for $1 million and run it with AI — ChrisGPT · 2026-07-22
- Netflix buys AI startup InterPositive for $587 million in cash — aloncarmel · 2026-07-22
- “Build what agents want” may become AI’s most crowded and commoditized category — vaibhavbetter · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22