Nathan Lambert says his RLHF book is finished after two years of nights and weekends
TheZachMueller · x · 2026-07-21
Nathan Lambert says his book Reinforcement Learning from Human Feedback is done.
He describes it as the resource he wished he had while learning to fine-tune, align, and now post-train models after ChatGPT. The book has been built over nights and weekends since 2024, and he says he tried to transfer as much of the intuition from building Olmo into book form as possible.
The post frames the book as a practical foundation for people working on RLHF and post-training rather than as a polished marketing launch.
Related event: Nat Lambert Completes RLHF Book with 10+ Hour Course(11 posts)→
More from Companies & People
- IIT Madras Launches EdTech Tulna Standards for AI-Powered Learning Products — ravi_iitm · 2026-09-11
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- X drama: Anthropic researchers accused of spying on academic customers and racing them to results — basedjensen · 2026-09-11
- Investor argues Palantir-Nvidia partnership should slash Anthropic's IPO valuation — pdamodaran · 2026-09-11