OLMo author says their RLHF book is finished after nights and weekends since 2024
giffmana · x · 2026-07-23
OLMo author says their RLHF book is finished
The quoted post announces that the author’s book, Reinforcement Learning from Human Feedback, is done.
They describe it as the book they wish they had while learning to fine-tune, align, and now post-train models after ChatGPT. The book has been assembled through nights and weekends since 2024, and the author says they tried to carry over as much of the OLMo team’s practical intuition as possible.
The reply is just a congratulatory note, but the substance is the book launch itself.
Related event: Nat Lambert Completes RLHF Book with 10+ Hour Course(11 posts)→
More from Companies & People
- IIT Madras Launches EdTech Tulna Standards for AI-Powered Learning Products — ravi_iitm · 2026-09-11
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- X drama: Anthropic researchers accused of spying on academic customers and racing them to results — basedjensen · 2026-09-11
- Investor argues Palantir-Nvidia partnership should slash Anthropic's IPO valuation — pdamodaran · 2026-09-11