OLMo author says their RLHF book is finished after nights and weekends since 2024
giffmana · x · 2026-07-23
OLMo author says their RLHF book is finished
The quoted post announces that the author’s book, Reinforcement Learning from Human Feedback, is done.
They describe it as the book they wish they had while learning to fine-tune, align, and now post-train models after ChatGPT. The book has been assembled through nights and weekends since 2024, and the author says they tried to carry over as much of the OLMo team’s practical intuition as possible.
The reply is just a congratulatory note, but the substance is the book launch itself.
Related event: Nat Lambert finishes his RLHF book(11 posts)→
More from Companies & People
- Stanford Students Are Leaving for Startups, the Post Says — pmddomingos · 2026-07-23
- Anthropic Exec Doubts Major Model Labs Can Capture Most AI Industry Profits — teortaxesTex · 2026-07-23
- If It Still Needs FDEs, It’s Not AGI Yet — pmddomingos · 2026-07-23
- Company says it won’t chase super-app status and is focused on AGI, not B-end revenue — teortaxesTex · 2026-07-23
- Moonshot AI is accused of distilling Anthropic’s Fable for its K3 model — sull · 2026-07-23
- Brockman says OpenAI may reuse Sora tech in “secret efforts” while reshaping its leadership — shiringhaffary · 2026-07-23