Nat Lambert says his RLHF book is finished after nights-and-weekends writing
deliprao · x · 2026-07-21
A new RLHF book from Nat Lambert is finished and headed for launch
The post says Nat Lambert has completed Reinforcement Learning from Human Feedback, a book he says he wished existed when he was learning to fine-tune, align, and post-train models after ChatGPT.
- He built it through nights and weekends starting in 2024.
- The book is meant to capture the intuitions behind fine-tuning, alignment, and post-training.
- He says it distills lessons from building Olmo into book form.
Related event: Nat Lambert Completes RLHF Book with 10+ Hour Course(11 posts)→
More from Companies & People
- IIT Madras Launches EdTech Tulna Standards for AI-Powered Learning Products — ravi_iitm · 2026-09-11
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- X drama: Anthropic researchers accused of spying on academic customers and racing them to results — basedjensen · 2026-09-11
- Investor argues Palantir-Nvidia partnership should slash Anthropic's IPO valuation — pdamodaran · 2026-09-11