RLHF book finished after nights and weekends of work since 2024

vjkaruna · x · 2026-07-22

The author says their book, Reinforcement Learning from Human Feedback, is finished and headed to launch.

They describe it as the book they wished they had while learning to fine-tune, align, and post-train models after ChatGPT, and say it was built from nights and weekends since 2024 by translating lessons from building Olmo into book form.

Related event: Nat Lambert Completes RLHF Book with 10+ Hour Course(11 posts)→

Original post →

More from Companies & People

Companies & People channel →