Nathan Lambert says his RLHF book is finished after a year of nights and weekends
alexisjross · x · 2026-07-21
Nathan Lambert says his book Reinforcement Learning from Human Feedback is finished.
He describes it as the book he wished he had while learning to fine-tune, align, and post-train models after ChatGPT, and says it was built from night-and-weekend study and documentation work since 2024.
Related event: Nat Lambert Completes RLHF Book with 10+ Hour Course(11 posts)→
More from Companies & People
- IIT Madras Launches EdTech Tulna Standards for AI-Powered Learning Products — ravi_iitm · 2026-09-11
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- X drama: Anthropic researchers accused of spying on academic customers and racing them to results — basedjensen · 2026-09-11
- Investor argues Palantir-Nvidia partnership should slash Anthropic's IPO valuation — pdamodaran · 2026-09-11