OLMo author says their RLHF book is finished after nights and weekends since 2024

giffmana · x · 2026-07-23

OLMo author says their RLHF book is finished

The quoted post announces that the author’s book, Reinforcement Learning from Human Feedback, is done.

They describe it as the book they wish they had while learning to fine-tune, align, and now post-train models after ChatGPT. The book has been assembled through nights and weekends since 2024, and the author says they tried to carry over as much of the OLMo team’s practical intuition as possible.

The reply is just a congratulatory note, but the substance is the book launch itself.

Related event: Nat Lambert finishes his RLHF book(11 posts)→

Original post →

More from Companies & People

Companies & People channel →